Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
152,898 characters · 19 sections · 98 citation commands
Optimal Targeting in Fundraising: A Causal Machine-Learning Approach
\thispagestyle{empty}
Ineffective fundraising lowers the resources charities can use to provide goods. We combine a field experiment and a causal machine-learning approach to increase a charity's fundraising effectiveness. The approach optimally targets a fundraising instrument to individuals whose expected donations exceed solicitation costs. Our results demonstrate that machine-learning-based optimal targeting allows the charity to substantially increase donations net of fundraising costs relative to uniform benchmarks in which either everybody or no one receives the gift. To that end, it (a) should direct its fundraising efforts to a subset of past donors and (b) never address individuals who were previously asked but never donated. Further, we show that the benefits of machine-learning-based optimal targeting even materialize when the charity only exploits publicly available geospatial information or applies the estimated optimal targeting rule to later fundraising campaigns conducted in similar samples. We conclude that charities not engaging in optimal targeting waste significant resources.
JEL codes: C93; D64; H41; L31; C21
Keywords: charitable giving, optimal policy learning,individualized treatment rules
Fundraising is a costly activity: The 25 largest US charities spend between $5\%$ and $25\%$ of total donations on fundraising expenses And2011. These numbers are a matter of concern for two reasons. First, high fundraising costs leave a smaller proportion of overall donations to finance charitable projects. This can lead to an underprovision of the goods and services that charities provide and may, thus, lower welfare if the donors' utility depends on provision levels Rose1982, Name2013. Second, high fundraising costs also matter from the charity's perspective because donors are averse to financing overhead costs Tin2007, Gne2014. Hence, charities with excessive fundraising expenses will be less successful in raising donations. In conclusion, reducing disproportional fundraising costs can be crucial, from both the welfare and charity-management perspectives. However, while there is broad literature studying how different fundraising instruments such as matching grants and unconditional gifts affect donors' behavior And2013, previous research has paid less attention to how charities can increase the cost-effectiveness of fundraising.
In this paper, we exploit a causal machine-learning-based approach to maximize a charity's fundraising effectiveness: optimal targeting of fundraising activities to potential donors. \footnote{See ath19b for a review of causal machine-learning methods.} Machine-learning-based optimal targeting exploits the possibility that, due to heterogeneity in donors' preferences, the effects of any fundraising campaign are likely heterogeneous across individuals with different observable characteristics. \footnote{For example, donation motives (such as altruism, warm glow, and reciprocity) are heterogeneously distributed across individuals falk2018. Consequently, responses to fundraising activities that leverage (some of) these motives are likely heterogeneous as well.} Charities that ignore this heterogeneity may engage in loss-leading fundraising, for example, by directing a costly fundraising instrument to notorious nondonors or donors who, in response to the instrument, do not increase their donations enough to cover the instrument's cost. Against this backdrop, our key contribution is to use machine learning to identify a targeting rule (out of all feasible rules) that is optimal in the sense that it maximizes expected profits by avoiding loss-leading solicitations. \footnote{A targeting rule is feasible if it is a deterministic function of the observable characteristics.} Consequently, charities that engage in this type of optimal targeting can increase net donations raised (i.e., donations net of fundraising costs) and, hence, provide more goods and services.
We demonstrate the potential of machine-learning-based optimal targeting using a common fundraising instrument as an example: small unconditional gifts that accompany a solicitation letter. In theory, such unconditional gifts should increase donations by triggering a reciprocal reaction in recipients Fal2007. In practice, the evidence on the effectiveness of unconditional gifts is, however, mixed. Some studies suggest that unconditional gifts are an effective fundraising instrument, while others find that they do not affect average donations or even backfire and lower giving (see the literature section for a discussion). This type of effect heterogeneity might very well operate through individual heterogeneities, such as heterogeneous social preferences. Thus, fundraising gifts offer a promising context to study the benefits of machine-learning-based optimal targeting.
The goal of optimal targeting lies in directing a fundraising instrument (a gift in our example) to a subset of individuals such that the fundraising campaign's expected profits are maximized (i.e., the additional expected net donations collected by the campaign). Yet, who should charities target for that purpose? By definition, the optimal targeting rule states that the profit-maximizing targets are the so-called net donors: individuals whose expected additional donation is higher than the marginal fundraising cost. In an ideal world, the charity would know each individual's donation with and without the fundraising instrument (i.e., the gift) and could, thus, determine the set of net donors. In reality, however, this set is unknown to the charity. The reason is that donations under both conditions (gift vs. no gift) are unknown before the campaign. Moreover, even after the campaign, the charity could still only learn a donor's behavior under her assigned condition. These complications motivate our approach to identify optimal targets for fundraising activities.
Our approach to optimal targeting relies on two ingredients. First, it exploits random variation in the assignment of an unconditional gift at the donor level. To induce this variation, we teamed up with a charity that sends out solicitation letters to its donor base once a year. In this setting, we randomly assigned almost $20{,}000$ potential donors to a gift treatment (in which individuals received the solicitation letter with a gift) and a control group (in which they received only the letter). Second, as previously highlighted, our approach relies on machine-learning algorithms. Particularly, we let an algorithm learn the relationship between the individuals' observable characteristics and their expected donation behaviors in the experiment's treatment and control groups (i.e., with and without the gift). From this relationship, we then estimate the (out-of-sample) set of predicted net donors who should be targeted. This procedure establishes what we label the estimated optimal targeting rule. Importantly, although we cannot trace out the causal effect of the characteristics that drive the heterogeneity, the policy-relevant increase in net donations achieved by the targeting rule is identified.
More specifically, our paper draws on the following machine-learning methods. We implement the optimal-policy-learning algorithm of ath21, which extends the empirical welfare-maximization approach of kit18 by machine learning. In our main specifications, we then estimate optimal targeting rules with the exact policy-learning tree of zho2019 and show the sensitivity of our results to various alternative estimators (logit, logit lasso, classification and regression trees, and classification forest). Regarding data, we feed our algorithm with information from various sources and of different types, including socioeconomic characteristics, past donation data, and geospatial information. The geospatial information consists of publicly available information from Google Maps on economic and cultural facilities close to the potential donor's place of residence.
Our analysis yields two sets of main results. The first set concerns the warm list (i.e., the sample of previous givers). For this sample, we demonstrate that machine-learning-based optimal targeting substantially boosts the charity's net donations compared to benchmarks where everybody receives the gift (increase relative to this benchmark: $13.8\%$) or no one receives the gift (increase: $14.3\%$). We also show that these positive effects on net donations do not merely reflect pull-forward effects (i.e., shifts of donations from a later year to the experimental year). Next, we illustrate that the estimated optimal targeting rule is stable across two identical consecutive fundraising campaigns conducted in similar samples. To that end, we repeated our field experiment in a consecutive year with a new sample of warm-list individuals. We then highlight that, relative to our benchmarks, the charity could substantially increase net donations in the second year if it distributed gifts according to the optimal targeting rule estimated with first-year data. Another insight of the analysis is that our fundraiser can reap the full benefits of machine-learning-based optimal targeting by relying on data that should be easily accessible by all charities. For the gains of our approach to materialize, knowledge about who is in the warm list and publicly available geospatial information is sufficient. Consequently, we expect our approach to be widely applicable.
The second set of results concerns the cold list (i.e., the sample of previous nondonors). In contrast to the warm list, machine-learning-based optimal targeting in the cold list does not increase net donations compared to the no-gift benchmark. Furthermore, the estimated optimal targeting rule does not broaden the donor base enough to justify the gift's additional fundraising costs. We conclude that, in our context, the charity should not target the gift to cold-list individuals at all. Whereas the details of our findings may well be specific to the setting, a general conclusion is that charities that do not optimally target their fundraising efforts waste significant resources.
The paper is organized as follows. Section (ref) outlines our contributions to the literature, Section (ref) explains the institutional background and design of the experiment, and Section (ref) describes our empirical strategy. Section (ref) discusses our results, and Section (ref) concludes. Online Appendices (ref)--(ref) provide supplementary materials.
In the following, we detail how our study relates and contributes to various literature strands in economics and marketing.
\addvspace{0.1cm}\noindentEconomics literature. The first relevant strand of literature, the theoretical fundraising literature in public economics, provides a theoretical underpinning for why fundraising targeting can be beneficial. The argument is as follows. Many charities have properties similar to privately provided public goods And2013: contributions are voluntary and the provided goods are nonexcludable and nonrivalrous. In such a context, fundraising tools can counteract free-riding and, hence, underprovision problems And1988, Mor2000, Vest2003, And2003, And2013. \footnote{Which tool is suited to increase provision depends on the context. Solicitation letters (perhaps providing return envelopes) counteract the underprovision problem in settings with transaction costs And2003. By contrast, leadership gifts oppose the underprovision of threshold public goods And1988 and public goods under imperfect information Vest2003. Lotteries, in contrast, can increase efficiency in standard public goods settings Mor2000.} However, if fundraising is costly and charities must compete for donors, competition can push the costs to such high levels that the total service provision falls Rose1982, Ald2010, Ald2014. Targeting of fundraising instruments can then be a tool to maximize donations net of fundraising costs and, thus, provision levels. Specifically, Name2013 show that a charity's optimal strategy is to target net donors. However, for charities, it is challenging to follow this theoretical rule, as they cannot easily identify these optimal targets. Along these lines, our contribution to this literature is to explore how a data-driven machine-learning approach can help charities predict the set of net donors. Moreover, we study the external validity of estimated optimal targeting rules across fundraising campaigns for the first time.
A second emerging literature strand closely related to our work empirically studies how to target fundraising among heterogeneous donors. We are aware of only two papers that examine this topic. First, Adena2019 apply a targeting strategy to matching gifts, a fundraising tool where funds collected before a campaign top up donations above a threshold. Their main result is that charities can crowd in donations by conditioning these thresholds on past giving behavior (i.e., the thresholds are targeted). Second, Drou2021 focus on belief-based targeting. They conclude that charities can increase net donations by targeting information treatments to potential donors who hold incorrect (low) beliefs about others' donations. Our paper differs from these studies in two dimensions: First, Adena2019 and Drou2021 both build on conceptual considerations to identify dimensions of heterogeneity used for targeting. Following ath21, our paper takes a more agnostic and more flexible, data-driven approach. Particularly, our machine-learning algorithms autonomously uncover the most influential dimensions of heterogeneity based on observable donor characteristics. Consequently, the resulting estimated targeting rules can flexibly account for many different sources of heterogeneity (e.g., preferences or income), as long as they correlate with individuals' observable characteristics. Second, instead of considering debiasing or threshold matching, we focus on the unconditional gift as a different fundraising tool.
A third relevant literature strand is the literature on fundraising gifts List2012. While Fal2007 documents that unconditional gifts are a cost-effective tool to increase donations in the warm list, other studies paint a more scattered picture. For example, lan10 find zero effects of gifts in the warm list, and ALP2008 conclude that conditional on giving, gifts even lower contributions. yin2020coins present similar results. \footnote{Thank-you gifts that charities hand out after a donation also seem to lower subsequent donations NEW2012. Eck2018 directly compare unconditional and thank-you gifts and show that donors are twice as likely to give when they receive a high-quality unconditional gift.} We add to the looming discussion on the causes of effect heterogeneity by showing that individual heterogeneity alone is powerful enough to account for gifts' negative and positive effects. \footnote{Of course, differences in the overall setting (such as varying causes of charities) might also explain why different studies come to varying conclusions.} In particular, in our setting, some groups of potential donors increase donations in response to gifts, and other groups reduce their donations. This finding persists if we restrict our sample to the warm list only. While these findings are insightful in themselves, our contribution to this literature is to demonstrate that charities can effectively exploit the effect heterogeneities to target gifts optimally.
By focusing on effect heterogeneity, we contribute to a fourth literature strand, heterogeneous responses to fundraising that identifies five forms of heterogeneity: (a) characteristics of the charitable organization and the purpose of the charity okt00, vri15, (b) characteristics of the donors andr03,andr01,raj09,wie13, (c) donation motives or preferences of donors eck12,har07,KIZ2018, (d) past donation behavior schl97,has14, and (e) crowding out mee17. Instead of studying single, selected dimensions of heterogeneity, we combine a range of individual characteristics and past donation behavior. Additionally, our paper adds a rarely used determinant of heterogeneity to the analysis, geospatial characteristics. As this information is easily accessible to charities, it is particularly beneficial for optimal targeting. \footnote{don19 and glae18,glae20 show that geospatial characteristics are proxies for income and socioeconomic characteristics. One reason is neighborhood segregation heb20.}
\addvspace{0.1cm}\noindentMarketing literature. Our paper also relates to two strands of literature on the targeting of marketing interventions. The first strand applies machine-learning techniques to target marketing in contexts other than charitable giving. For example, GUE2015 and ASC2018 study targeting of retention efforts to lower customer churn in telecommunications, professional memberships, and insurance. Moreover, FON2019, ELL2020, GUB2020, and smi2021 study targeting of recommendations, promotions, coupons, and pricing (mainly in online marketing). These studies not only consider contexts very different from ours, but also employ other methods. Particularly, the papers estimate heterogeneous effects of (randomized) marketing interventions and transform them into a targeting rule by discretizing. \footnote{In marketing, this procedure is called “uplift modeling. ”Economists use the label “effect-based targeting.”} Our paper, instead, follows ath21 and estimates the optimal targeting rule directly (instead of first estimating the effect heterogeneity). To the best of our knowledge, our paper is the first application of this novel, more direct method, explicitly tailored for optimal targeting. \footnote{Related machine-learning-based literature profiles or segments customers based on their responses under one condition. For example, CUI2006 and KIM2005 show that neural networks are valuable tools to improve targeting in response models. Moreover, ABE2004 successfully apply reinforcement learning to cross-channel marketing, and SCH2017 apply a multiarmed bandit to target display advertising.} The second relevant literature strand is the literature on charitable marketing WIN2013, KIZ2018. While this literature has studied targeting in charitable giving, it so far has not used the powerful machine-learning-targeting toolkit. To sum up, the available studies deviate from our paper either because they use machine learning in contexts other than charitable giving or because they study targeting in charitable giving without using causal machine-learning techniques.
\addvspace{0.1cm}\noindentTargeting literature. Methodologically, we contribute to the small but rapidly growing literature that applies machine-learning methods to target public and private policies and18,hit18,kan13,kna18b,kni19,roc11,kle15. While these papers consider such contexts as taxation and labor-market programs, our study is the first that applies a fully fledged optimal policy learning algorithm like that of ath21 to the context of charitable giving.
Our approach to derive a machine-learning-based optimal targeting rule proceeds in three steps. First, we conduct a field experiment that randomly allocates our fundraising instrument, an unconditional gift. Second, we use machine-learning algorithms to estimate the optimal targeting rule for this gift in a random subsample of the experimental data while retaining the remaining sample. Third, we extrapolate the estimated optimal targeting rule to the retained sample and apply off-policy-learning techniques to assess the estimated rule's out-of-sample performance. While this section details the experimental design and data, Section (ref) outlines the machine-learning and off-policy learning approaches.
In $2014$, we implemented a natural field experiment in collaboration with a fund\-raiser of the Catholic Church that operates in a German urban area. \footnote{In $2015$, we implemented a second experiment with two treatments, an unconditional gift treatment and a gift treatment that framed the gift as a reward for past donations. We use this experiment to (a) examine the stability of the targeting rules and (b) evaluate and compare the effects of the differently framed gifts in a companion paper.} For decades, this fundraiser has organized a large-scale, annual fundraising campaign: Once a year, it has mailed solicitation letters to all resident church members, irrespective of previous donations. This fund drive aims to finance local church-related projects, such as the renovation of clergy houses, parish centers, or churches. Our experiment exploited this campaign by (a) experimentally altering how the fundraiser contacted potential donors in $2014$ and (b) analyzing individuals' behavior in $2014$ and $2015$.
\addvspace{0.1cm}\noindentControl group. Individuals in our experiment's control group received the standard solicitation letter, the contents of which remained unchanged from the pre-experiment years. Particularly, the letter highlighted the fundraiser's cause and asked recipients for a donation. To lower transaction costs, the fundraiser distributed the solicitation letter together with a remittance slip, prefilled with the fundraiser's bank account and the donor's name. In the pre-experiment years, potential donors received identical transaction forms. \footnote{For years, donations in the context of the fund drive could be made exclusively via bank transfer.}
\addvspace{0.1cm}\noindentGift treatment. Our design of the gift treatment closely follows Fal2007. In $2014$, individuals in this treatment received the solicitation letter together with an unconditional gift. The gift consisted of three envelopes paired with different folded cards, picturing Albrecht D\"{u}rer's “immaculate flower studies” (see Figure (ref)). Further, we added one sentence to the solicitation letter, stating that the fundraiser “would like to provide the included folded cards as a gift.” The total per-unit cost for mailing the control-group solicitation letter amounted to 0.43 euro (printing plus postage). In the gift treatment, the per-unit cost increased by $1.16$ euro (postcards and envelopes: $0.47$ euro; boxing and additional postage: $0.69$ euro). From $2015$ onward, all individuals in the sample received a solicitation letter very similar to the one distributed in the pre-experiment years.
\addvspace{0.1cm}\noindentSample. In $2014$, $26\%$ of the urban area's population were members of the Catholic Church. We drew a sample from this population consisting of $2{,}354$ warm-list individuals (individuals who had donated at least once before the experiment) and $17{,}425$ cold-list individuals (individuals who had never donated before). These individuals were then randomly allocated to the control and treatment groups, exploiting a stratified randomization scheme. Particularly, we assigned $1{,}180$ of the warm-list individuals to the gift treatment and $1{,}174$ to the control group. \footnote{The strata were defined based on list (warm vs. cold), gender, household type indicators, quintiles of individuals' predicted baseline willingness to give, and quintiles of age. To construct a proxy for the baseline willingness to give in the treatment year, we first regressed an indicator variable for giving in the year before the experiment on indicator variables for further lags of the giving indicator. We then used the estimated model to predict the probability of giving in the treatment year (out of sample).} By contrast, $2{,}283$ cold-list individuals received the treatment, while $15{,}142$ were part of the control group.
\addvspace{0.1cm}\noindentThe setting's benefits. Our setting serves as a suitable testing ground for machine-learning-based optimal targeting. First, as the fundraiser did not employ any targeting strategies before the experiment, the setting offers a clean environment to study our machine-learning approach's potential. Second, it provides rich data that not only allow us to estimate powerful targeting rules, but also enable us to test which type of data are especially beneficial for machine-learning-based optimal targeting (see the following description of the data). Third, because the fundraiser contacts all church members exhaustively, the setting offers the possibility to study the cold and warm lists separately. We, hence, can not only examine the optimal targeting of gifts among past donors, but also explore whom to target in the process of acquiring new donors. Fourth, because we were able to gather data for two postexperiment years, the setting allows us to study if the estimated optimal targeting rule increases total donations or pulls forward donations from $2015$ to $2014$. Also, we can examine the rule's external validity. Fifth, because religious giving dominates the charitable giving landscape List2011, targeting is particularly relevant in this context. \footnote{For example, in Germany, church-related causes benefit the most from private giving: They receive approximately $35\%$ of total private giving Sepnd2016. No other type of cause benefits from a similarly high share of total donations. The numbers for the United States are very similar And2013.}
\noindentData sources. Our study draws on two separate, comprehensive data sets. The first set of data includes administrative records provided by the Catholic Church. The records hold a number of socioeconomic characteristics, such as gender, marital status, and age. Furthermore, they contain individual-specific information on donations for the years $2006$--$2015$. Accordingly, we observe all potential donors' donation histories for eight pre-experiment years and their donations in the first two years after the experiment. Our second data source is Google Maps. Specifically, we used the Google Maps API to collect geospatial information on economic and cultural facilities near each individual's residence. We then merged this data with the administrative records based on postal addresses. In particular, we collected the number of restaurants, supermarkets, medical facilities, cultural facilities, and churches within $300$ meters of each the home address. We also web-scraped the distance from the home address to the central train station, city hall, main church, and airport. Furthermore, we retrieved the elevation of the home address. The main reason for using these geospatial characteristics is that they are readily available to charities and are powerful proxies for income and other socioeconomic characteristics don19, glae18,glae20 that likely explain response heterogeneity. Taken together, our algorithms for optimal targeting rely on three types of input data: (a) socioeconomic information, (b) information on past donation behavior, and (c) publicly available geospatial information. Note that charities that manage to collect even more comprehensive datasets could further improve machine-learning-based optimal targeting.
\addvspace{0.1cm}\noindentDescriptive statistics. Table (ref) in Online Appendix (ref) reports the descriptive statistics for the donation amount and the donation probability. Furthermore, supplementary tables either study the balance of observable characteristics across the warm and cold lists (see Table (ref) in Online Appendix (ref)) or the control and treatment groups (see Tables (ref) and (ref)). Several features of the data stand out. First, unsurprisingly, cold-list individuals donated much less in the first year after the experiment (average donation: $0.18$ euro) compared to warm-list individuals (average: $16.02$ euro). Their donation probability is also much lower. Similar results emerge when summing up the donations made in the first two years after the experiment. Second, donations are highly right-skewed and have excess kurtosis, highlighting that few donors give very large gifts. The following analysis highlights that, despite this data feature complicating our prediction task, we are nevertheless able to estimate effective optimal targeting rules. Third, cold list and warm list individuals differ in socioeconomic characteristics (see Table (ref)). \footnote{On average, cold list individuals are younger, have a higher likelihood of being single, and tend to have a shorter residency duration in the urban area. Individuals in the warm list donated an average of four times, with a total donation amount of $126$ euro over the eight years before the experiment. By construction, cold list individuals have a donation history of zero donations. Individuals in the cold list live, on average, closer to the city center (closer to city hall and the central station) than individuals on the warm list. Close to the home address (within $300$ meters), cold-list individuals have, on average, more access to restaurants, supermarkets, medical and cultural facilities, and churches than warm-list individuals.} This observation suggests that individual characteristics might explain heterogeneous giving behavior. Fourth, the observable characteristics are balanced across the treatment and control groups (see Tables (ref) and (ref)).
This section describes the estimation and identification strategy of machine-lear\-ning-based optimal targeting rules and how to assess the rules' out-of-sample performance. We start by introducing conditional average treatment effects (called CATEs) in Subsection (ref). The CATEs formally describe heterogeneous treatment effects as a function of observable characteristics. In Subsection (ref), we then introduce binary optimal targeting rules that are related to the continuous CATEs. These rules maximize the charity's net donations by assigning individuals to either the targeted group or the untargeted group. Recall that our approach directly estimates the optimal targeting rule. This feature distinguishes it from approaches that first estimate the continuous CATEs and then derive the rules as nonlinear transformations of these effects. In the final step, we discuss our classification approach in Subsection (ref). In Subsection (ref), we further discuss how we measure our targeting rules' out-of-sample performance compared to benchmark rules.
\noindentNotation. We use the potential-outcome framework rub74 to describe the parameters of interest. The potential-outcome framework is useful when studying targeting because it allows us to describe an individual's reaction under different (counterfactual) treatment conditions. The treatment variable $D_i$ indicates whether a fundraising gift was sent to individual $i$ (for $i = 1, \ldots, N$), with
$Y_{i}(1)$ denotes the potential donations in response to the solicitation letter with a fundraising gift. $Y_{i}(-1)$ denotes the potential donations in response to the letter without the gift.
\addvspace{0.1cm}\noindentCausal effects. Using the previous notation, the individual causal effects are
In an ideal world, the charity would know $\delta_{i}$. It could then (exclusively) assign the gift to individuals for whom $\delta_{i}$ exceeds the gift's cost. However, because $Y_{i}(1)$ and $Y_{i}(-1)$ cannot be observed simultaneously, the fundamental problem of causal analysis is that $\delta_{i}$ is unobservable. Nevertheless, it is possible to identify and estimate group averages of $\delta_{i}$. For example, the average treatment effect (ATE), $\delta = E[\delta_i]= E[Y_{i}(1)-Y_{i}(-1)]$, is the expected average effect of the gift on donations. Moreover, there might be effect heterogeneity with regard to observable characteristics, $X_{i}$, which allows researchers to identify even finer-grained subgroup-specific effects. For example, andr01 and andr03 show that men and women differ in their donation behavior. In this vein, the CATE describes the association between an individual's characteristics $x$ and the expected effect of sending the fundraising gift to the individual:
\addvspace{0.1cm}\noindentIdentification of causal effects. The ATEs and CATEs are identified under the stratified experimental design and the stable unit treatment value assumption (SUTVA):
(see proof in Online Appendix (ref)). The proof distinguishes between strata characteristics, which we call $Z_i$, and heterogeneity variables, $X_i$. Due to stratified randomization, we need to account for the strata characteristics to achieve identification. In contrast, we do not require the heterogeneity variables for identification. They, however, are potentially associated with heterogeneous effects of the gift.
\noindentOptimal targeting rule. A targeting rule, $\pi(X_i) \in \{-1, 1\}$, is a deterministic function that assigns the gift to prospective donors based on their observable characteristics, $X_i$. Under the rule, individuals with $\pi(X_i)=1$ receive the solicitation letter with the gift, and individuals with $\pi(X_i)=-1$ receive the solicitation letter without the gift. The purpose of the optimal targeting rule is to maximize the expected net donation $P(\cdot)$ of the fundraising campaign, defined as the expected donation minus the gift's variable costs. Formally, the expected net donation is
where $Y_i(\pi(X_i))$ is the donation amount of individual $i$ under the rule $\pi(X_i)$ and $c$ are the variable costs of the gift. We ignore fixed costs, as they do not alter the targeting rule.
\addvspace{0.1cm}\noindentBenchmarks rules. To evaluate the gains of optimal targeting, we compare the expected net donations under the optimized rule $P(\pi(X_i))$ to the expected net donations under three benchmarks: a rule that assigns the gift to everybody (all-gift benchmark), a rule that assigns the gift to no one (no-gift benchmark), and a rule with random allocation (random-gift benchmark). First, we can contrast the optimal targeting rule to the all-gift benchmark $\pi(X_i)=\pi_1=1$. For this benchmark, the expected net donation is $P(\pi_1)= E[Y_i(1)]- c$. Consequently, the excess net donation of the optimal targeting rule compared to this benchmark is
Second, equivalently, we compare the optimal rule to the no-gift benchmark $\pi(X_i)=\pi_{-1}=-1$, under which the expected donations are $P(\pi_{-1})= E[Y_i(-1)]$. Thus, relative to this benchmark, optimal targeting increases net donations by
Third, we consider the random-gift benchmark, $\pi_R$, under which each individual has a $50\%$ probability of receiving the gift. Given that $\pi_R$ triggers the expected net donation of $P(\pi_R)= 1/2 \cdot (E[Y_i(1) + Y_i(-1)]- c)$, the excess net donation of the optimal rule is
The random rule can be viewed as a default option when no information about the effectiveness of the fundraising instrument is available, and the fundraiser has no preferences about the allocation of the instrument.
Two further points are of note. First, the optimal targeting rule which maximizes the net donations $P(\cdot)$, also maximizes $Q_{1}(\cdot)$, $Q_{-1}(\cdot)$, and $Q_{R}(\cdot)$. The reason is that $P(\pi_{1})$, $P(\pi_{-1})$, and $P(\pi_{R})$ are constant. Second, for the estimation of the optimal targeting rule, we maximize the sample analog of (ref). Because we do not observe the individual causal effects, $\delta_i$, which we need to determine (ref), we first discuss how to approximate these parameters.
To introduce our estimation strategy of the optimal targeting rule, we proceed in two steps. In the first step, we discuss augmented inverse probability weighting (AIPW) to estimate an approximation of the individual causal effects. In the second step, we show how to use the AIPW score to estimate the optimal targeting rule.
\addvspace{0.1cm}\noindentAugmented inverse probability weighting. An essential ingredient for the optimal targeting rule is $\delta_i$. As we mentioned before, $\delta_i$ is unobservable and cannot be estimated directly. However, an approximation score of $\delta_i$ can be sufficient to estimate the optimal targeting rule. The AIPW score,
which depends on the experimental strata characteristics $Z_i$, is an example of such an approximation score. \footnote{Alternatively, kit18 suggest inverse probability weighting scores, and bey09 propose offset weighting scores.} The so-called nuisance parameters are the conditional expectations of the donations, ${\mu}_{1}(z) = {E}[Y_i |D_i=1,Z_i=z]$ and ${\mu}_{-1}(z) = {E}[Y_i |D_i=-1,Z_i=z]$, and the conditional probability that the gift was sent ${p}(z) = {Pr}(D_i= 1|Z_i=z)$. The latter is often called the propensity score. Under the SUTVA and the experimental design, the expected value of the AIPW score identifies the ATE $\delta = E[\Gamma_i]$. \footnote{Note that the nuisance parameters, ${\mu}_{1}(z) = {E}[Y_i |D_i=1,Z_i=z]= {E}[Y_i(1) |Z_i=z]$ and ${\mu}_{-1}(z) = {E}[Y_i |D_i=-1,Z_i=z]= {E}[Y_i(-1) |Z_i=z]$, equal conditional expectations of the potential donations with and without the gift under the SUTVA and the experimental design.} The conditional expectations of the AIPW score identify the CATEs, $\delta(x)= E[{\Gamma}_i|X_i=x]$. For completeness, we sketch the identification proofs for the AIPW score in Online Appendix (ref).
\addvspace{0.1cm}\noindentEstimating the AIPW score. We can estimate the AIPW score as follows. First, we estimate the nuisance parameters. In this step, we obtain the estimated conditional expectations of the potential donations with and without the gift by $\hat{\mu}_{1}(z)$ and $\hat{\mu}_{-1}(z)$ and the estimated propensity score by $\hat{p}(z)$. Second, we plug the estimated nuisance parameters into the estimator of the AIPW score, $\hat{\Gamma}_i= \hat{\Gamma}_i(1)-\hat{\Gamma}_i(-1)$, with
and
The corresponding average-treatment-effect estimator
is consistent, asymptotically normal, and semiparametrically efficient under the requirement that the nuisance parameter estimators are consistent and converge sufficiently fast Chernozhukov2017,Robins1994. In our application, we have precise information about the stratification process. Therefore, we use parametric nuisance parameter estimators which satisfy the requirements. In particular, we use a logit to estimate the propensity score and OLS to estimate the conditional expectations of the potential donations with and without the gift (Table (ref) in Online Appendix (ref) reports the estimated coefficients of the different models). \footnote{Online Appendix (ref) presents two robustness tests regarding the specification of the nuisance parameters. First, it replaces the estimated propensity score with the population propensity score. Second, it estimates all nuisance parameters with cross-fitted logit-lasso models instead of conventional estimators. The results are quantitatively identical to our baseline results (see Tables (ref) and (ref) in Online Appendix (ref)).}
Note that we could alternatively estimate the ATEs with a OLS regression model. However, in contrast to OLS, the AIPW estimator does not impose any linearity assumptions and does not restrict effect heterogeneity. The latter is particularly relevant for the targeting approach. Having said this, we show in Section (ref) that the OLS and AIPW estimates of the ATEs are similar.
\addvspace{0.1cm}\noindentEstimating the optimal targeting rule. For the estimation of the optimal targeting rule, ath21 propose replacing the unobservable individual causal effect in the sample analog of (ref), $\delta_i$, with the estimated AIPW score, $\hat{\Gamma}_i$:
Alternatively, the objective function (ref) can be formulated as the weighted classification estimator
where $(\hat{\Gamma}_i-c)=\mbox{sign}(\hat{\Gamma}_i-c) \cdot |\hat{\Gamma}_i - c |$ bey09,zadr03,zhao12. This estimator aims to classify the sign of the net donation effects and weigh each observation by $|\hat{\Gamma}_i - c |$. The objective function is maximized when the signs of $\pi(X_i)$ and $(\hat{\Gamma}_i-c)$ are equal. If some signs differ, misclassifications of individuals who respond strongly (i.e., individuals with large weights) reduce the net donations more than misclassification of individuals who do not respond strongly to the gift (i.e., individuals with small weights). Accordingly, the optimal targeting estimator should prioritize individuals with large weights. Because the estimated optimal targeting rule is a deterministic function of the observable characteristics $X_i$, there is an implicit connection between optimal targeting and the CATEs, even though we estimate the optimal targeting rule directly (without estimating the CATEs first).
The main result of ath21 that enables estimation of (ref) is that, when the complexity of the estimator is restricted, the optimal targeting rule $\pi^*$ achieves asymptotically minimax-optimal regret man04. Along these lines, in principle, any restricted weighted classification estimator could be used to estimate (ref). We, however, follow ath21 and use shallow decision trees to estimate the optimal rule. Trees partition the sample into mutually exclusive strata based on the heterogeneity characteristics, $X_i$. Furthermore, the tree depth restricts the complexity of the estimated optimal targeting rule, which makes trees suitable estimators in our context. In our main specifications, we follow zho2019 and use exact policy-learning trees, with a search depth of two, to estimate the optimal targeting rule. \footnote{For implementation, we use the R package policytree sver20.}
\addvspace{0.1cm}\noindentAdvantages of decision trees. It is possible to use standard estimators, such as a weighted logit regression, to estimate (ref). However, decision trees have several advantages compared to logit regressions for the estimation of optimal targeting rules. They select the relevant heterogeneity characteristics in a data-driven way by balancing the bias-variance trade-off. This feature is particularly useful when we have no a priori domain knowledge about the relevant characteristics. Even if we know the relevant characteristics, there might be several highly correlated measures of these characteristics, and it may be a priori unclear which are the most relevant. For example, in our application, several geospatial characteristics are highly correlated, and there is little a priori guidance on which should be used. In the extreme case, including too many highly correlated characteristics in a logit regression could cause multicollinearity problems. Furthermore, it is typically unclear how flexible the empirical model should be with regard to nonlinear and interaction terms. Trees can automatically incorporate nonlinear and interaction terms of the different characteristics without precoding. This feature minimizes the risk of overlooking important heterogeneities. Having said this, we study the sensitivity of our results to different estimation methods for the targeting rule in Section (ref). In particular, we consider logit, logit lasso, classification and regression trees (CART), and classification-forest estimators.
\addvspace{0.1cm}\noindentEstimating the gains of targeting. Once we have estimated the optimal targeting rule, $\pi^*$, we can apply the sample analogy principle to estimate the gains of targeting relative to the benchmarks:
These estimators are consistent, asymptotically normal, and semiparametrically efficient unify18.
We estimate the gains of targeting using a cross-validation procedure. Our procedure randomly partitions our data into $K=20$ equally sized samples. We then use $K-1$ partitions to estimate the targeting rule $\pi^*$ and calculate $\hat{P}(\pi^*(X_i))$, $\hat{Q}_1(\pi^*(X_i))$, $\hat{Q}_{-1}(\pi^*(X_i))$, and $\hat{Q}_R(\pi^*(X_i))$ in the retained partition. We repeat this procedure, discarding each of the $K$ partitions once. In this way, we use the entire dataset efficiently. Finally, we report the average values of $\hat{P}(\pi^*(X_i))$, $\hat{Q}_1(\pi^*(X_i))$, $\hat{Q}_{-1}(\pi^*(X_i))$, and $\hat{Q}_R(\pi^*(X_i))$ over all $20$ partitions. The cross-validation approach allows us to assess the estimated targeting rules' out-of-sample performance. It also addresses the concern that the targeting rule reflects spurious relationships and overstates the success of targeting due to overfitting.
This section presents our results. Subsection (ref) discusses the ATEs of the gift on donations, and Subsection (ref) explores the heterogeneity of the effects. The section proceeds by discussing the effectiveness of our optimal targeting approach in Subsection (ref) before describing several properties of the estimated optimal targeting rule in Subsection (ref). Subsection (ref) outlines which characteristics are sufficient to increase profits significantly through machine-learning-based optimal targeting. Finally, Subsection (ref) examines the stability of the estimated optimal targeting rule, and Subsection (ref) explores the robustness of our results to alternative estimators.
To facilitate comparison to the literature studying the effects of unconditional gifts on donations, our first step is to estimate the ATE of the gift on donations. Table (ref) reports three different estimates, focusing on behavior in the first year after the experiment: estimates from unconditional OLS regressions (Columns 1 and 4), estimates from conditional OLS regressions (Columns 2 and 5), and estimates from the previously introduced AIPW estimator (Columns 3 and 6). \footnote{In contrast to the OLS regressions, the AIPW estimator relies on fewer functional-form assumptions, allows for heterogeneous treatment effects, and is more robust to misspecification.} Columns 1--3 cover the warm list and Columns 4--6 the cold list.
\addvspace{0.1cm}\noindentAverage treatment effects for the warm list. Two observations characterize the responses in the warm list. First, the gift increased average donations by $1.21$--$1.24$ euro, though the effects are not statistically significant (see Row A in Table (ref)). \footnote{This result is in line with lan10, who also report insignificant effects for the warm list.} Notably, these estimates do not account for the cost of the gift (which were $1.16$ euro). Second, when accounting for the costs, we find small and insignificant net-of-cost effects between $0.05$ euro and $0.08$ euro, depending on the chosen estimator (see Row B in Table (ref)). The small values imply that we neither find evidence for the hypothesis that the gift treatment was, on average, profitable (i.e., increased average donations by more than the costs) nor that it resulted in a net loss for the fundraiser. The two observations are insightful from a targeting perspective. To see why, note that we can think of the treatment effects as driven by a change from the benchmark targeting rule where no one receives the gift (control group) to the one where everybody receives the gift (treatment group). \footnote{More formally, the expected difference in net donations between the benchmark rules $\pi_{1}$ and $\pi_{-1}$ is $P(\pi_{1})-P(\pi_{-1})=E[\delta_i]-c$.} Along these lines, the insignificant net-of-cost effects speak against the hypothesis that the all-gift benchmark outperforms the no-gift benchmark in terms of available funds.
\addvspace{0.1cm}\noindentAverage treatment effects for the cold list. Very different results emerge for the cold list. The ATEs are significant (note the larger sample size), but much smaller, and amount to just $0.19$ euro (Row A). As a result, when accounting for costs, the all-gift benchmark rule would result in a significant loss of $0.97$ euro per donor, compared to the no-gift benchmark rule (Row B). Accordingly, considering only these two benchmark rules, we find that the more profitable strategy is to send the gift to no one in the cold list. \footnote{The previous literature has reported similar results in the past. For example, ALP2008 show that gifts increase donations, but the increase is insufficient to cover the gift's costs.}
Taken together, we conclude that naive targeting of gifts to all individuals in the warm and cold list does not increase our fundraiser's net donations. Such a strategy would even likely result in a net loss as there are many more cold-list than warm-list individuals. Based on these insights, one might be tempted to conclude that our fundraiser can never increase net donations with the gift. In the following, we, instead, show that a more targeted gift campaign can be very successful.
This subsection provides a descriptive analysis suggesting that the charity can likely increase raised net donations by deviating from the no-gift and all-gift benchmarks. For that purpose, we demonstrate that the effects of the gift are heterogeneous. To motivate the analysis, note that in the absence of heterogeneous treatment effects, the charity cannot benefit from targeting rules that are more flexible than the benchmarks. Intuitively, as all individuals respond similarly to the gift, the optimal targeting rule would either correspond to the no-gift or all-gift benchmark. However, with effect heterogeneity, some individuals may increase their donations by more and others by less than the gift's cost. If this is the case, the charity's optimal strategy would be to deviate from the benchmarks and target a subset of individuals only.
\addvspace{0.1cm}\noindentSorted-effects approach. As a first step, we use the CATE-based sorted-effects approach of sort to study descriptively if there is sufficient heterogeneity for optimal targeting. This method allows us to visualize the distribution of the effects of the gift on donations while reporting confidence intervals that account for multiple testing. Particularly, we specify a linear OLS regression including all observable characteristics plus interactions with the treatment dummy. Using this model, we estimate the CATE for each individual and report the percentiles of the estimated CATEs (labeled sorted effects). \footnote{Statistical inference is based on a multiplier bootstrap sort.} We then examine if the sorted-effects model shows heterogeneous effects below and above the gift's costs to investigate if deviations from the two benchmarks are likely beneficial. Note, however, that we only use the sorted-effects approach to assess the potential of optimal targeting descriptively (i.e., we do not use it to identify targets). Instead, we estimate the optimal targets using a machine-learning approach explicitly developed for this purpose.
\addvspace{0.1cm}\noindentHeterogeneity in the warm list. Panel (a) in Figure (ref) depicts the heterogeneity of the treatment effect on the donation amount for the warm list (solid line). It sorts the estimated CATEs by size and plots the size of the treatment effect in euro (vertical axis) against the percentiles of the effect size (horizontal axis). The red horizontal line represents the cost of the gift (1.16 euro).
The figure reveals substantial treatment-effect heterogeneity. \footnote{One might be interested in whether the size of the treatment effects correlates with observable characteristics. Tables (ref) and (ref) in Online Appendix (ref) report the mean values of all characteristics for the groups with the $10\%$ largest and the $10\%$ smallest sorted effects. In the warm list, the individuals with the $10\%$ largest effects tend to have donated less before the experiment and to live at a lower altitude than individuals with the $10\%$ smallest effects, although the effects are insignificant. In the cold list, the individuals with the 10% largest effects tend to live significantly closer to the city center than individuals with the 10% smallest effects.} For some individuals, the treatment effects are positive, which is in line with the sequential-reciprocity hypothesis Duf2004,Fal2007. For example, the $5\%$ most responsive individuals increase donations by more than $17.99$ euro in response to the gift. By contrast, at the fifth percentile, donations decrease by $11.59$ euro. Such an adverse effect points to the possibility that even a “warm” gift (folded cards) may change the donors' perception of the relationship with the fundraiser from a communal to an exchange norm yin2020coins. The pronounced heterogeneity is interesting for at least two reasons. First, and most importantly, the heterogeneity indicates that machine-learning-based optimal targeting can be highly beneficial in the warm list: For $45\%$ of all individuals, the estimated effects exceed the cost of providing the gift. Thus, by targeting these individuals and not targeting donors with effects lower than the cost, the charity could substantially increase net donations. Second, the figure also reveals that the individual heterogeneity alone is powerful enough to account for the negative and positive effects reported in the literature.
\addvspace{0.1cm}\noindentHeterogeneity in the cold list. Panel (b) in Figure (ref) highlights that the effect heterogeneity in the cold list is much smaller than in the warm list. For example, the donation amount decreases by $0.73$ euro at the fifth percentile and increases by $1.38$ euro at the $95$ percentile. Moreover, we find that the gift-induced increase in donations exceeds the costs only for individuals above the $92$ percentile in the cold list, and even for these individuals the size of the treatment effects is relatively small.
\addvspace{0.1cm}\noindentHeterogeneity test. Rather than only presenting descriptive analyses, we also more formally test if there is detectable effect heterogeneity based on observed characteristics. For this purpose, we implement the best linear predictor method of Chernozhukov2020. Particularly, following ath19b, we use a test that relies on a causal forest estimator. Table (ref) in Online Appendix (ref) documents the findings. \footnote{Positive and statistically significant coefficients of the heterogeneity loadings provide evidence for detectable effect heterogeneity based on the observed characteristics. The causal forest approximates effect heterogeneity well when the coefficients of the heterogeneity loadings are close to one.} In line with the previous results, the test detects significant effect heterogeneity in the warm list and no heterogeneity in the cold list. We conclude that the potential to increase net donations by subgroup-specific targeting is much lower in the cold list than in the warm list.
This subsection evaluates if machine-learning-based optimal targeting allows us to exploit the documented response heterogeneity to increase net donations in the first year after the experiment. We first estimate optimal targeting rules ath21. We then evaluate the effectiveness of the estimated rules in raising net donations by comparing their out-of-sample performance to our benchmarks. For that purpose, we use the cross-validation approach described in Section (ref).
\addvspace{0.1cm}\noindentThe estimated optimal targeting rule in the warm list. Table (ref) focuses on the warm list and documents how the estimated optimal targeting rule performs out of sample relative to the benchmarks. The table provides three sets of insights. First, Panel A reports that the estimated targeting rule recommends deviating from the no-gift and all-gift benchmarks. Specifically, it assigns the gift to $33\%$ of the warm-list individuals (Column 1).
Second, Panel B documents that, compared to our benchmarks, the charity would benefit substantially from applying the estimated optimal targeting rule. In the first year after the experiment, the average net donation under the estimated optimal targeting rule is $17.61$ euro (Column 1). This value implies that, under the estimated optimal targeting rule, the average donation, net of costs, is $2.14$ euro ($13.8\%$) higher than if everybody received the gift (Column 2), $2.20$ euro ($14.3\%$) higher than if no one received the gift (Column 3), and $2.17$ euro ($14.1\%$) higher than if the gift was randomly allocated to one half of the warm-list sample (Column 4). Accordingly, the estimated optimal targeting rule is significantly more profitable than all three benchmark policies. In conclusion, our techniques allow the fundraiser to increase net donations significantly and, hence, service and goods provision.
Third, Panel C documents that, by applying the estimated optimal targeting rule, the fundraiser would also impact outcomes besides net donations (labeled secondary outcomes). Thus, although we train our algorithm to maximize net donations, implementing the estimated rule would trigger secondary effects as a byproduct. One secondary outcome is the donation probability within the first postexperiment year (see Row C1). Specifically, by implementing the estimated optimal rule, the fundraiser would increase this probability by almost three percentage points ($5\%$) compared to the no-gift benchmark. Our proposed targeting strategy, hence, not only maximizes net donations, but also broadens the donor base compared to a scenario without gifts. In contrast to this result, the donation probability under the estimated rule is not significantly higher than under the all-gift benchmark. Hence, although this benchmark endows many more individuals with the gift, it does not fundamentally increase the donation probability. This result suggests that intensive margin responses drive the difference in net donations between the benchmark and the estimated rule. Besides impacts on the donation probability, the table also reveals secondary effects on longer-term outcomes. For example, Row C2 of Table (ref) demonstrates that, when using the outcome “aggregate donations made within two years after the experiment,” the positive effects of applying the estimated optimal targeting rule persists. This is an important result from the fundraiser's perspective: It highlights that the machine-learning-induced increase in net donations would not curb subsequent donations.
\addvspace{0.1cm}\noindentThe estimated optimal targeting rule in the cold list. Table (ref) shows the out-of-sample performance of the estimated optimal targeting rule in the cold list. Due to the substantial size of the cold-list sample ($17{,}000$ individuals), all of the effects are very precisely estimated. Again, the results are very different from those for the warm list. One marked difference is that the estimated target group is much narrower in the cold list (Panel A): In line with the evidence from the sorted-effects model, the estimated optimal rule assigns the gift to just $1.4\%$ of the cold-list individuals. Given this finding, it is not surprising that the average net donation under the estimated optimal rule ($0.15$ euro) is virtually identical to that under the no-gift benchmark (Column 3 in Panel B). By contrast, the estimated optimal targeting rule outperforms the all-gift (Column 2 in Panel B) and random-gift benchmarks (Column 4 in Panel B). The reason is that campaigns that apply these two benchmarks would result in losses (due to the gifts' costs). Table (ref) also presents evidence on secondary effects (Panel C). Again, we find no evidence for pull-forward or delay effects. Further, there are only minimal impacts on the donation probability. To sum up, the potential of machine-learning-based optimal targeting in the cold list is very limited, perhaps because the cold list consists of many notorious nondonors. This insight was not clear a priori and could only be established with a flexible, data-driven approach such as ours.
A common theme in the fundraising literature is understanding the characteristics of individuals who give to charitable causes And2013. Our machine-learning approach allows us to extend this literature by performing a broader descriptive analysis: Instead of merely describing the characteristics of givers, we can differentiate between the predicted net donors and predicted net recipients (i.e., individuals who increase donation by less than the gift's cost). The analysis, hence, reveals the characteristics of the individuals who, according to our estimated targeting rule, should receive the gift and contrasts them with the characteristics of those who should not be targeted.
\addvspace{0.1cm}\noindentCharacteristics of predicted net donors and net receivers in the warm list. Table (ref) focuses on the warm list. It reports the means and standard deviations of all observed characteristics for predicted net donors (Columns 1--2) and predicted net receivers (Columns 3--4). It also shows the standardized difference, a standard balance diagnostic (Column 5). The table highlights some apparent differences between the two groups. For example, before the experiment, the predicted net donors donated on average more and also more frequently than the predicted net recipients (Panel A). They also live in more central areas that are characterized by (a) lower altitudes and (b) a more lively environment with more churches, restaurants, and cultural facilities (Panel B). \footnote{In the urban area we study, the topology correlates with distance to the city center. Concretely, individuals living in lower altitudes live on average closer to the city hall and the main church.} By contrast, the differences in socioeconomic characteristics are not very pronounced. If anything, predicted net donors are more frequently female (Panel C). When interpreting these results, keep in mind that the characteristics are not necessarily reflecting channels through which the impacts of the gift operate. The patterns might instead mirror effects working through correlated unobservable variables. For example, the donation history may approximate individuals' general willingness to give, and geospatial information could proxy income. In this vein, these and similar unobservable variables might channel the responses to the gift. We consider it a strength of our approach that the machine-learning algorithm can pick up the underlying forces that shape individuals' reactions to the gift without the need to collect data or explicitly model the relationships.
\addvspace{0.1cm}\noindentCharacteristics of predicted net donors and net receivers in the cold list. Table (ref) reports similar results for the cold list. The predicted net donors from the cold list also live in more central areas than the respective predicted net recipients, but the areas now tend to be less lively. Regarding the socioeconomic characteristics, there are more pronounced differences compared to the warm list. Relative to predicted net recipients, predicted net donors tend to have a higher likelihood of being females, singles, and newly settled residents.
\addvspace{0.1cm}\noindentDecision trees. A second, natural way to describe the estimated targeting rule is to plot decision trees. However, because we use a cross-validation approach to evaluate the estimated rule's out-of-sample performance, we obtain $20$ different trees per list. To reduce complexity, Figure (ref) in Online Appendix (ref), hence, reports trees estimated using the entire sample. The first tree concerns the warm list. It shows that the previous donation amount and the elevation at the home address serve as split variables. By contrast, the second tree for the cold-list is split based on the number of restaurants near the residence and the distance to the city hall and airport. Importantly, the trees do not identify the causal effect of the donor characteristics on the net donations. Rather, they indicate correlations between donor characteristics and the gift's causal effect. Hence, we must interpret them cautiously.
As discussed before, we estimate the optimal targeting rule using socioeconomic characteristics, donation history, and geospatial information. The next step of our analysis explores which of these data are especially powerful to estimate targeting rules. It also investigates if our algorithms require all the data to estimate effective optimal targeting rules. Besides being interesting in itself, studying this topic is vital for charities. Data collection is costly, and, frequently, some forms of data (such as socioeconomic characteristics) are unavailable. In many settings, a charity might only have access to address data before sending out written solicitations. Hence, from a charity's perspective, the usefulness and feasibility of machine-learning-based optimal targeting critically depend on the data required to target net donors effectively.
\addvspace{0.1cm}\noindentRelevant data in the warm list. Table (ref) focuses on the warm list and explores the relevance of the different data types for the performance of the estimated optimal targeting rule. For this purpose, it evaluates the performance of several estimated targeting rules, each using only a subset of the available data. As before, the table benchmarks these more sparsely estimated rules against the all-gift (Column 3), no-gift (Column 4), and random-gift (Column 5) benchmarks. Additionally, Column 6 compares these newly estimated rules to the estimated optimal (baseline) rule obtained when using all data (reported in Table (ref)).
The results are as follows: First, our baseline rule clearly outperforms a rule that relies only on socioeconomic characteristics. Also, the rule based only on socioeconomic characteristics does not outperform the three benchmarks (Panel A). As socioeconomic data are often hard to obtain, these results might not be problematic for charities. Second, our baseline rule does not significantly dominate rules that either use only data on past donations (Panel B) or use only geospatial information (Panel C). Similar to our baseline rule, these two more sparsely estimated rules also significantly outperform the three benchmarks. These findings suggest that the data on past donations and geospatial information are substitutes in targeting. Hence, charities that do not have access to details on individuals' donation histories might instead rely on publicly available geospatial data only. Third, Panels D--F further emphasize that the socioeconomic characteristics are of very limited use for machine-learning-based optimal targeting. Adding them does not improve the rules that rely solely on the donation history (Panels B and D) or geospatial information (Panels C and F). Furthermore, a rule that combines the donation history with the geospatial information performs as well as a rule that uses all the data. The reason is that our baseline rule is not relying on the socioeconomic characteristics at all.
\addvspace{0.1cm}\noindentRelevant data in the cold list. Table (ref) reports similar analyses for the cold list. To that end, it restricts the data either to socioeconomic characteristics or geospatial information. It turns out that the socioeconomic characteristics are also redundant in the cold list: Again, our baseline targeting rule does not use socioeconomic characteristics. Thus, the results do not change when using only the geospatial information instead of all data (see Panel B).
We draw two main conclusions from this subsection. First, in the warm list, the fundraiser can significantly improve its campaigns' profits by relying only on widely available geospatial information. Put differently, the fundraiser does not necessarily need access to detailed data on past donations, and socioeconomic characteristics seem to be of little use for optimal targeting. This finding raises the attractiveness of our approach for charities that, for the sake of simplicity or due to data-collection costs, prefer to rely on a single data source. Second, in the cold list, the potential benefits of machine-learning-based optimal targeting are very limited. Given the data available, the dominant strategy is not to send the gift to individuals in the cold list.
This subsection explores if the estimated optimal targeting rule is externally valid across two identical consecutive fundraising campaigns. To that end, we test whether the charity can increase net donations in the second follow-up campaign by distributing gifts according to the rule estimated on first-campaign data. We consider two dimensions of external validity, which we label “time stability” and “sample stability.” According to our definition, an estimated optimal targeting rule is time stable if the charity can beneficially apply it to a follow-up campaign run on a similar sample of potential donors. Instead, a rule that satisfies sample stability can even be successfully applied to consecutive campaigns run on samples with different characteristics.
\addvspace{0.1cm}\noindentFollow-up experiment. We study the estimated rule's external validity by exploiting data from a follow-up experiment that took place in a subsequent year using different participants. Specifically, in $2015$, we randomly allocated a new sample of $3,616$ warm-list individuals to the control group and the gift treatment (treatment probability: $50\%$). \footnote{Given the evidence from the $2014$ campaign, we decided to consider only warm-list individuals. In Table (ref) of Online Appendix (ref), we report the balance of the observed characteristics by treatment status.} Notably, the $2015$ sample vastly differs from the $2014$ sample in its observable characteristics. For example, individuals in the $2015$ sample donated more, and more frequently before the experiment, than those in the $2014$ sample (see Table (ref) in Online Appendix (ref)). Their response to the gift was also more pronounced: The average effect of the gift on net donations was $1.33$ euro in $2015$ compared to $0.08$ in $2014$.
\addvspace{0.1cm}\noindentStudying both dimensions of external validity. The follow-up experiment allows us to study both dimensions of external validity. Because the $2015$ and $2014$ samples differ, we can test sample stability by examining how the rule estimated on $2014$ data (hereinafter, $2014$ rule) performs in the full $2015$ sample. A test of time stability, instead, requires similar samples in both years. To construct comparable samples, we employ a caliper propensity score matching approach with a radius of $0.1$ that matches the $2015$ sample to the $2014$ sample (see Appendix (ref) for details). We then evaluate time stability by assessing the $2014$ rule's performance in the matched $2015$ sample.
\addvspace{0.1cm}\noindentResults of our external-validity checks. Table (ref) demonstrates how net donations would have changed in $2015$ when the charity would have allocated the gift following the $2014$ rule. Again, it presents results relative to our three benchmarks. Furthermore, it separately considers the matched sample and the full $2015$ sample. The results are as follows: First, the $2014$ rule suggests that $35.6\%$ of the warm-list individuals in the matched sample and $35.3\%$ in the full $2015$ sample should receive the gift. Second, the $2014$ rule seems to be time stable. The charity would have benefited substantially from applying the $2014$ rule to the matched $2015$ sample. Hereby, it could have increased net donations by $2.64$ euro relative to the all-gift benchmark (Column 2 in Row B1) and $2.30$ euro relative to the no-gift benchmark (Column 3 in Row B1). These effects are comparable to those in $2014$ (see Table (ref)). Third, however, we cannot confirm sample stability when considering the full $2015$ sample. By applying the $2014$ rule to this sample, the charity would have been unable to increase net donations relative to the all-gift and no-gift benchmarks. \footnote{The coefficient for the all-gift benchmark is negative because, in $2015$, the gift significantly increased average net donations. Hence, it is difficult to outperform the all-gift benchmark.} This result is not surprising: If the sample changes too much from $2014$ to $2015$, the second-year data contains few individuals similar to those on which the algorithm was trained. This likely affects the performance of the optimal targeting rule in $2015$, at least in the presence of effect heterogeneity (individuals in the $2014$ and $2015$ samples respond differently).
\addvspace{0.1cm}\noindentVarying the sample differences. Given the previous discussion, one occurring question is how much the charity can change the sample such that the optimal targeting rule is still effective. Online Appendix (ref) provides evidence on this topic. Specifically, it explores how similar the $2014$ and $2015$ samples must be for the $2014$ rule to still increase net donations in $2015$. To that end, we sequentially increase the caliper radius, resulting in the $2015$ sample differing more and more strongly from the $2014$ sample. Figure (ref) in Online Appendix (ref) presents the results. One message of the figure is that the $2014$ rule significantly outperforms the no-gift benchmark up to a caliper radius of $0.25$ (i.e., for quite different samples). By contrast, because the gift increases average net donations in the full sample, it is much more difficult to outperform the all-gift benchmark. For caliper radii above $0.1$, the effects are no longer statistically significant (but still positive in expectation).
To sum up, in our context, the estimated optimal targeting rule is time stable over one year; it is also stable over different samples, but only if they do not differ too much. Thus, when using previously estimated rules, our charity needs to hold both samples as similar as possible to obtain the best result (e.g., by using an appropriate sampling design). Even though this finding is not surprising, it never has been shown before.
Our final step is to investigate the robustness of the estimated optimal targeting rules to different estimation approaches. First, we consider alternative depths of the exact policy-learning tree (one and three). \footnote{In the cold list, exact policy-learning trees of “depth three” are infeasible due to computational constraints.} Second, we compare the results of the exact policy-learning trees to the results of standard CARTs brei84. Third, for the CARTs, we also consider one version in which we use cross-validation to select the tree depth in a data-driven way. \footnote{We use the Gini index for tree splitting and a $10$-fold cross-validation procedure to select the optimal tree depth. In the warm list, the number of terminal leaves varies between four and thirteen across the 20 different cross-validated trees, with an average of $5$. In the cold list, the number of terminal leaves varies between one and nineteen across the $20$ different cross-validated trees, with an average of $3.6$.} Fourth, we employ weighted logit as a standard estimator. Fifth, we use two alternative machine-learning estimators: logit lasso hast16 and classification forests brei01. \footnote{We build $1{,}000$ trees for the classification forest. We draw a $50\%$ random subsample with replacement for each tree and randomly select $50\%$ of the baseline characteristics. We use the Gini index for tree splitting. We restrict the minimum size of the terminal leaf to $50$ observations.} For the logit estimator, we consider two different model specifications. The baseline specification includes all observed characteristics linearly ($24$ variables in the warm list and $16$ variables in the cold list). The flexible specification additionally includes squared terms of continuous variables and first-order interactions between most characteristics ($320$ variables in the warm list and $148$ variables in the cold list). The logit-lasso method selects the relevant characteristics from the flexible model specification. \footnote{We specify the penalty $\lambda$ of the logit lasso that minimizes the misclassification error using a $10$-fold cross-validation approach.}
Table (ref) in Online Appendix (ref) reports the results of the sensitivity analysis for the warm list, and Table (ref) focuses on the cold list. In the warm list, our baseline optimal targeting strategy clearly dominates the alternative specifications: The baseline optimal targeting rules of all the alternative estimators yield lower net donations than our main specification. Specifically, the exact policy-learning tree with depth one and the logit with the flexible model specification have the lowest out-of-sample performance. The CART with the cross-validated tree depth is the only alternative estimator that also significantly outperforms the all-gift and no-gift benchmark allocation rules. For the cold list, the logit-lasso specification and the CART with cross-validated tree depth yield $0.01$ euro higher net donations than our main specification. However, our main finding that the fundraiser should not send the gift to cold-list individuals persists.
This paper studies machine-learning-based optimal targeting of fundraising instruments by exploiting data from a natural field experiment. The underlying idea of optimal targeting is that fundraisers can maximize a campaign's profits by directing a fundraising instrument to individuals who increase their donations in response to the instrument by more than its cost. We label those individuals net donors. However, charities do not observe the set of net donors. We employ a machine-learning algorithm to estimate the relationship between individual characteristics and the potential donors' response to small unconditional gifts. Based on this algorithm, we can predict the subset of net donors and, hence, estimate machine-learning-based optimal targeting rules.
Our paper's first key message is that, in the warm list, machine-learning-based optimal targeting substantially boosts the charity's net donations. In our application, net donations increase by about $14\%$ compared to the all-gift and no-gift benchmarks. Notably, the benefits of machine-learning-based optimal targeting even materialize when (a) applying the estimated optimal targeting rule to a later campaign conducted in a similar sample or (b) relying only on widely available geospatial data. Hence, charities can easily apply the proposed strategies to raise additional funds, net of costs. The second message is that, in the cold list, the approach does not raise donations sufficiently to cover the fundraising instrument's additional costs. We conclude that the fundraiser should not target cold-list individuals at all. We also document that the increase in net donations stems from heterogeneities in donors' responses to fundraising activities. Previous literature has suggested that such heterogeneities exist, for example, by providing mixed evidence on the effectiveness of unconditional gifts on giving Fal2007,yin2020coins,ALP2008.
In conclusion, our paper demonstrates that machine-learning-based optimal targeting can significantly increase the cost effectiveness of fundraising. One particularly noteworthy benefit of our applied machine-learning toolkit is that it allows charities to target fundraising efforts agnostically in a wide variety of contexts. Thus, to optimize targeting, charities do not need to develop a theoretical foundation or make strong assumptions on the functional relationship between individual characteristics and giving. We are, therefore, confident that the proposed approach offers an accessible way forward to improve the effectiveness of fundraising. We are also looking forward to evolving research applying similar techniques to alternative fundraising instruments and different settings. It will also be interesting to see how our findings generalize across these alternative tools and environments.
\setcounter{page}{1} \setcounter{footnote}{0} \setcounter{equation}{0} \setcounter{footnote}{0}