EconBase
← Back to paper

The Privacy-Utility Trade-Off of Location Tracking in Ad Personalization

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

127,673 characters · 34 sections · 57 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

The Privacy–Utility Trade-Off of Location Tracking in Ad Personalization

abstractFirms collect vast amounts of behavioral and geographical data on individuals. While behavioral data captures an individual’s digital footprint, geographical data reflects their physical footprint. Given the significant privacy risks associated with combining these data sources, it is crucial to understand their respective value and whether they act as complements or substitutes in achieving firms’ business objectives. In this paper, we combine economic theory, machine learning, and causal inference to quantify the value of geographical data, the extent to which behavioral data can substitute it, and the mechanisms through which it benefits firms. Using data from a leading in-app advertising platform in a large Asian country, we document that geographical data is most valuable in the early “cold-start” stage, when behavioral histories are limited. In this stage, geographical data complements behavioral data, improving targeting performance by almost 20%. As users accumulate richer behavioral histories, however, the role of geographical data shifts: it becomes largely substitutable, as behavioral data alone captures the relevant heterogeneity. These results highlight a central privacy–utility trade-off in ad personalization and inform managerial decisions about when location tracking creates value. {\bf Keywords:} Privacy, Personalization, Advertising, Informational Complementarity, Policy Evaluation, Spatial Statistics

\thispagestyle{empty}

Introduction

Mobile devices dominate the digital economy, reshaping how consumers interact with media, commerce, and advertising. Nearly all Americans own a mobile phone, and nine in ten use a smartphone pew2024smartphone, which creates an ideal environment for targeted marketing. In response, mobile advertising now accounts for more than two-thirds of all digital ad spending in the United States, surpassing \$200 billion in 2024 emarketer2024mobile. Much of this growth stems from advances in tracking technologies that allow advertisers to monitor users’ digital and physical activities and tailor messages to individual consumers. While these capabilities have made mobile advertising a central channel for marketers, they have also heightened privacy concerns, raising questions about whether the economic value of such data practices justifies the associated privacy risks.

Privacy concerns arise most directly from two types of data collected by advertising platforms: behavioral data, such as click histories or in-app purchases, and geographical data, such as GPS coordinates or IP-based locations ghose2019mobile. Behavioral data captures a user’s digital footprint within apps, while geographical data captures their physical footprint—their movements and proximity in the real world. Geographical information is particularly sensitive, as even a few location points can be sufficient to re-identify individuals de2013unique. When combined with behavioral data, these risks amplify: the two data types complement each other in enabling more precise inference of personal attributes and increasing the likelihood of re-identification, even when datasets are anonymized de2018privacy.

Given the complementarity of geographical and behavioral information in amplifying privacy risks, it is crucial to examine whether a similar complementarity exists in the factor that motivates firms to continue collecting them—the economic value they create. In particular, it remains unclear whether behavioral and geographical data act as substitutes or complements in generating economic value for advertising platforms. Addressing this question requires first establishing how to measure the value of each data type, then comparing their relative contributions and interactions, and finally identifying the mechanisms through which geographical data may provide distinct informational value. These considerations motivate the following research questions:

enumerate• How can we measure the value of geographical data in ad personalization relative to behavioral data? • To what extent do geographical data generate value once behavioral data are available, and how does this value change as behavioral data accumulate? • Through which mechanisms do geographical data provide distinct information?

We face several challenges in answering these questions. The first challenge is conceptual: how should we define the value of information? A common approach focuses on prediction accuracy, for example, whether adding geographical data improves click-through forecasts. However, predictive gains alone do not guarantee better advertising outcomes ascarza2018retention. What ultimately matters is whether available information improves decision quality. Accordingly, we draw on economic theory to define the value of geographical and behavioral data through their impact on decision outcomes. This framework allows us to assess whether the two sources act as complements, where their joint use delivers gains beyond the sum of their individual effects, or as substitutes, where the contribution of one diminishes once the other is available. For example, if behavioral data indicate that a user is a New York Knicks fan, knowing that the user is in New York City is valuable for advertising tickets to a home game, so the two information sources may act as complements. However, for online Knicks merchandise sold online and shipped nationwide, location adds little value once behavioral interest is known, so the two sources may act as substitutes. This distinction is central, as collecting information beyond its marginal value imposes privacy costs without improving outcomes, whereas complementary information can justify broader data use.

The second challenge is empirical. Our objective is to assess whether geographical data provides incremental decision value beyond behavioral data. This requires a behavioral model that fully captures the dynamic structure of user histories; otherwise, estimated gains from geographical data may reflect behavioral misspecification rather than informational value. Because behavioral data are inherently temporal, we leverage a best-in-class Long Short-Term Memory (LSTM) network with an attention mechanism to capture temporal dependencies and heterogeneous relevance across past interactions.

The third challenge is statistical. Our objective is to test whether two sources of information act as substitutes or complements, which requires comparing policy values under alternative information sets. This, in turn, demands counterfactual outcomes that are not directly observed: we only observe user responses under the policies actually deployed, not under alternative targeting strategies. To address this limitation, we apply inverse propensity scoring (IPS) horvitz1952generalization, which recovers counterfactual policy values by reweighting observed outcomes using assignment probabilities. A key feature of our setting is the platform’s quasi-proportional auction mechanism, which is directly observed in the data and allocates ads in proportion to the quality-adjusted bid. This mechanism provides plausibly exogenous variation and supports reliable propensity score estimation and credible statistical inference.

Together, these challenges motivate a unified framework that combines economic theory to define information value through policy comparisons, machine learning to model user outcomes from complex behavioral and geographical data, and causal inference to enable counterfactual evaluation via IPS. We apply this framework to data from a leading mobile advertising network in a major Asian country, covering 10.3 million impressions from 439{,}157 users over a 10-day period. We consider four targeting scenarios: a benchmark \(X^{\varnothing}\) that uses only contextual information (e.g., app category or time of day); \(X^G\), which augments context with geographical data such as exact latitude and longitude; \(X^B\) augments contextual features with behavioral information from users’ prior impression and click histories; and \(X^{GB}\), which combines both geographical and behavioral information. Differences in data structure across scenarios motivate the use of best-in-class models, with gradient-boosted trees (XGBoost) for scenarios based on static and cross-sectional features (\(X^{\varnothing}\) and \(X^G\)) and LSTM networks with attention for scenarios that incorporate sequential behavioral histories (\(X^B\) and \(X^{GB}\)). Leveraging the observed quasi-proportional auction mechanism to support credible propensity score estimation, we apply IPS to recover counterfactual click-through rates (CTRs) and formally quantify the value of behavioral and geographical information, providing a statistical test of whether the two sources act as complements or substitutes.

At the aggregate level, we compare all targeting regimes to the baseline CTR. Targeting without user-level data (\(X^{\varnothing}\)) increases CTR by 20.44%. Adding geographical information (\(X^G\)) raises CTR by 28.7%, behavioral information (\(X^B\)) by 30.3%, and combining both (\(X^{GB}\)) yields an improvement of 41.5%. Although the joint regime achieves the highest performance, aggregate results alone do not provide systematic evidence that geographical and behavioral information act as complements or substitutes. This motivates examining heterogeneity as behavioral information accumulates.

As users receive more ad impressions, advertisers observe richer behavioral histories, which change both the value of behavioral information and its interaction with geographical information. In the earliest stage (1--2 impressions), behavioral information is minimal, and performance is driven largely by geographical information, with the combined regime performing similarly to geographical data alone. In the intermediate stage (approximately 5--25 impressions), behavioral information accumulates but remains sparse. In this range, geographical and behavioral information act as complements, and using both information sources yields gains beyond either source alone. Once users exceed 25 impressions, behavioral information becomes sufficiently rich to capture preferences independently, and geographical information becomes a substitute, adding little incremental value. Overall, these results show that geographical information evolves with impression depth: complementary when it is sparse, and substitutable once it is rich.

Our findings have direct implications for managers and policymakers. Geographical information creates value primarily in the short run, acting as a temporary complement when behavioral information is sparse, but becomes substitutable as user histories accumulate. For firms, this implies that geo-targeting can be useful during the cold-start phase but delivers little incremental benefit once behavioral models are established. For regulators, the results underscore a key trade-off: location information entails substantial privacy risks while offering limited long-term value in digital advertising. Firms should therefore reassess the strategic role of geographical information, using it selectively when behavioral histories are unavailable and reducing reliance on it as richer user profiles emerge.

In summary, our paper makes several contributions to the literature. First, while prior research has emphasized the value of either behavioral or geographical data in isolation, we provide a systematic comparison of the two. Second, we develop a unified framework that combines economic theory to define information value, machine learning to model advertising responses, and causal inference to evaluate counterfactual outcomes. This framework allows managers and policymakers to assess whether different information sources act as complements or substitutes, providing a practical tool to evaluate targeting decisions while accounting for privacy costs. Substantively, we show that both geographical and behavioral data create value, but their roles evolve as behavioral information accumulates: geographical data complements behavioral data when histories are sparse and becomes substitutable once behavioral information is rich. This pattern implies that the marginal value of geo-targeting is highest when behavioral information is sparse and declines as behavioral histories become richer, raising questions about the privacy–utility trade-off.

The remainder of the paper proceeds as follows. We review related work on the value of information, advertising and privacy, and spatial economics in \S(ref). We introduce the institutional setting of mobile advertising and describe the large-scale dataset in \S(ref). We then formalize the problem within a decision-under-uncertainty framework to define the value of information in \S(ref). Next, we first highlight key challenges and then present our empirical framework in \S(ref), where we define complementarity and substitutability and combine machine-learning–based click prediction with inverse propensity scoring for counterfactual evaluation. We report empirical findings on the value of behavioral and geographical data and their heterogeneity in \S(ref), and we examine the mechanisms behind these patterns using a residualized spatial autocorrelation test in \S(ref). We discuss implications for data use, privacy, and targeting efficiency in \S(ref), and we conclude by summarizing our contributions and outlining directions for future research in \S(ref).

Related Work

Our work relates to a broad literature on how the value of information is defined and measured for decision-making under uncertainty. Theoretically, blackwell1953comparison provides the benchmark by ordering signals according to whether they increase expected reward across all decision problems (the Blackwell order), and borgers2013signals extend this framework to multiple signals by characterizing when information sources act as complements or substitutes based on whether one signal’s marginal value rises or falls in the presence of another. Empirically, rossi1996value quantify the value of alternative information sets for direct marketing by comparing profits from targeted couponing based on purchase histories and demographic characteristics. Complementing this perspective, kim2022selecting show that the value extracted from rich data depends on the chosen level of data granularity and model specification, formalizing the bias–variance trade-off in information use. Subsequent work examines how data access and targeting policies affect outcomes using causal strategies. Using randomized experiments, ascarza2018retention compares targeting rules based on churn risk versus treatment-effect lift, cui2019learning manipulate the availability of information shown to consumers, and wernerfelt2025estimating experimentally restrict access to offsite behavioral data to estimate its incremental value. Other studies exploit natural experiments, such as aridor2024evaluating, who use Apple’s App Tracking Transparency as an exogenous shock to behavioral data access to identify its impact on advertising outcomes. Finally, off-policy evaluation methods recover policy values from logged or observational data, as in rafieian2023ai and rafieian2023optimizing. We extend this literature by proposing a unified framework that systematically compares multiple information sources and empirically tests whether they act as complements or substitutes through their marginal contributions to decision outcomes in high-dimensional settings.

Second, our study connects to the literature on privacy and personalized advertising, particularly in user tracking and engagement modeling. In display advertising, early work by goldfarb2011online shows that ad intrusiveness and privacy sensitivity significantly affect engagement, with well-targeted yet subtle ads performing best. Subsequent research demonstrates that personalized ads can increase relevance but also intensify privacy concerns: tucker2014social show that personalization in social networks raises demand for stronger privacy controls, while acquisti2016economics formalize the trade-offs firms face between personalization benefits and consumer resistance to data collection. More recent studies examine the consequences of restricting data access, with johnson2020consumer quantifying the revenue losses from consumer opt-outs and Rafieian2021 showing that privacy restrictions reduce targeting efficiency and may affect market competition. In this regard, despite well-documented privacy risks of location data de2013unique, relatively few studies directly examine the value of geographical information. Closely related to our work, narang2025privacy quantify the predictive value of geo-tracking data for forecasting consumer visits. We extend this line of research by moving beyond predictive accuracy and developing a unified framework for substitutability/complementarity between geographical and behavioral data that equips us to more precisely assess the privacy--utility trade-off.

Third, a line of empirical research in marketing and economics studies geographic dependence in outcomes using methods from spatial econometrics. Early contributions introduce diagnostics for spatial autocorrelation. moran1950notes proposes Moran’s \( I \) as a global measure of spatial dependence in outcomes, while geary1954contiguity develops Geary’s \( C \) to capture more localized spatial variation. Subsequent work applies these diagnostics to substantive economic settings. bronnenberg2001unobserved model spatial dependence in market shares and promotions across neighboring markets to account for correlated, unobserved retailer actions. Focusing on longer-run demand patterns, bronnenberg2009brand document persistent geographic structure in brand demand, showing that brands retain higher shares near their historical origins. Related research examines diffusion and contagion by exploiting geographic or network adjacency. A body of empirical research in marketing studies how spatial structure shapes consumer behavior and diffusion outcomes bradlow2005spatial. At finer levels of granularity, larson2005exploratory show that spatial layout affects consumer exposure and choice through physical shopping paths. This perspective extends to diffusion and contagion, where researchers exploit geographic or network adjacency: manchanda2008role use physician networks to separate targeted communication from peer influence, iyengar2011opinion study social networks to identify opinion leadership effects, and bollinger2012peer use zip-code–level exposure to quantify neighborhood spillovers in solar adoption. Despite growing evidence of substantial spatial heterogeneity in advertising effectiveness, as documented by luo2025mapping, existing work largely focuses on detecting and modeling spatial correlation in outcomes rather than evaluating geographical data as an input to ad personalization. We address this gap by assessing the value of geographical information for personalization and by proposing a residual spatial autocorrelation (RSA) framework that decomposes spatial correlation conditional on behavioral data to pin down the channel through which geographical information creates value.

Setting & Data

We start by describing the institutional setting of the mobile advertising platform (\S(ref)). Next, we detail the dataset, including impressions, clicks, and contextual variables (\S(ref)). We then explain our sampling strategy and report summary statistics (\S(ref)). Finally, we describe the train–test split design used to evaluate model performance (\S(ref)).

Setting

Our data are sourced from a leading mobile in-app advertising platform in a large Asian country, which held over 85% of the mobile advertising market during the time of our study. The platform serves as an intermediary between advertisers and mobile apps (publishers) and is responsible for delivering over 50 million ad impressions daily. This marketplace consists of four primary players:

itemize• Users are mobile app consumers who generate impressions and may choose to click on displayed ads. • Publishers are app developers that integrate ads into their apps and monetize based on ad clicks. • Advertisers design banner ads and specify per-click bids. They can target users based on variables such as province, smartphone brand, app category, internet service provider (ISP), hour of the day, and connectivity type. The ad network does not support detailed personalized targeting. • Platform or ad network manages real-time auctions to match impressions with ads. Ads are placed as bottom banners and are refreshed every minute. Only clicks result in payment under a cost-per-click (CPC) scheme.
figure[figure omitted — 209 chars of source]

Figure (ref) illustrates the structure of the in-app advertising marketplace. When a user opens an app, the platform initiates a real-time auction among eligible ads, allocating impressions through a quasi-proportional rule mirrokni2010quasi: ads with higher bid–quality products are assigned higher probabilities of being shown, but selection remains stochastic. For example, if two ads have bid–quality products of 2 and 1, the first ad is twice as likely to be shown as the second, but both ads retain strictly positive probabilities of exposure. As a result, propensity scores are non-zero by design, which ensures overlap across ads. Importantly, quality scores are fixed and not personalized at the user level, so allocation probabilities do not depend on unobserved user characteristics, implying unconfounded assignment by design. If the user remains active beyond one minute, a new impression is generated, and a new auction is run.

Data

Following the setting explained above, we have data on all impressions and click outcomes observed over a 30-day period from September 30, 2015, to October 30, 2015. During this period, we observe a total of 1,594,831,699 ad impressions and 14,373,293 clicks, resulting in an overall click-through rate (CTR) of approximately 0.90%. For each impression, the dataset includes detailed information across several dimensions.

For each impression, the dataset records a comprehensive set of variables. First, we observe (1) Timestamp, capturing the exact time of the impression, and (2) AAID, the Android Advertising ID, a user-resettable unique device identifier that enables anonymous tracking across applications. At the ad-delivery level, we record (3) AppID, the identifier of the app displaying the ad, and (4) CreativeID, the identifier of the specific ad creative shown. We also observe (5) Bid, the advertiser’s submitted bid amount (fixed throughout the sample period), and (6) CPC, the cost-per-click charged in the event of a click. Geographical attributes include \textit{(7) Latitude}, \textit{(8) Longitude}, and \textit{(9) Province}, which are available for most impressions. Further contextual information comprises \textit{(10) Connectivity}, indicating whether the user was on Wi-Fi or cellular data, \textit{(11) Brand}, the smartphone manufacturer, \textit{(12) MSP}, the mobile service provider, and \textit{(13) ISP}, the internet service provider. Finally, \textit{(14) Click} is a binary indicator equal to one if the user clicked on the ad and zero otherwise.

A key aspect of this dataset is that it is sourced directly from the ad platform and includes the complete set of variables that advertisers could potentially use for targeting. As such, we observe the same information available to both the platform and advertisers at the time of impression delivery. This mitigates common concerns in observational studies around hidden targeting mechanisms or unobserved confounding. In our context, the completeness of the data supports modeling assumptions such as conditional ignorability, which are crucial for later analyses involving counterfactual analysis and policy evaluation.

Sampling and Summary Statistics

To enable user-level analysis while maintaining computational tractability, we restrict attention to impressions from the top 10 ads served between October 20 and October 30, yielding an initial sample of 16{,}662{,}783 impressions. We address missing values, particularly in variables that serve as key features in our analysis, by excluding records without geographic information, specifically latitude and longitude. After this filtering, the working dataset contains 10{,}336{,}703 impressions and has no remaining missing values in any other columns. From this dataset, we identify new users who joined the platform during the first three days of the window, October 20 to October 22, and follow their subsequent activity through October 30. This design allows us to observe user responsiveness from the start of their platform interaction, yielding a cohort of 439,157 unique users.

We now present some summary statistics for key categorical variables in the dataset. Table (ref) reports, for each variable, the number of unique categories, the share of impressions associated with the top three values, and the number of non-missing observations.

table[table omitted — 936 chars of source]

We observe a total of 9,515 unique apps, with the top three accounting for a sizable share of total impressions. Among the ten ads included by design, exposure is uneven, with the most frequently shown ad representing over 22% of impressions. The distribution of user identifiers, based on Android Advertising IDs, is highly diffuse, with no single user accounting for more than 0.02% of total impressions. Other variables, such as smartphone brand, connectivity type, ISP, and province, exhibit varying degrees of concentration, reflecting both common patterns and localized variation in usage across the population.

While Table (ref) highlights variation in exposure across contextual features, it does not reveal how user responsiveness may differ across behavioral or geographical information. To explore this further, we present descriptive evidence on how CTR varies with user history and spatial location. These patterns help motivate the relevance of behavioral and geographical data for downstream modeling tasks.

Behavioral Heterogeneity

We present two descriptive results that illustrate how behavioral patterns shape user CTR. First, we examine how the length of a user’s exposure history relates to CTR. Second, we analyze the relationship between past click behavior and the likelihood of future clicks.

Figure (ref) plots the cumulative share of impressions against the cumulative share of clicks, where users are ordered by the number of prior impressions they have seen, so that impressions are ranked from early to late in a user’s exposure history, and each point on the curve compares the share of total impressions up to that history length with the share of total clicks they generate. The curve lies well above the $45^\circ$ line: a disproportionate share of clicks comes from early exposures (e.g., the first 25% of impressions, corresponding to short histories, generate nearly half of all clicks). Figure (ref) shows the probability of a click on the next impression, $\Pr(\text{click}_{t+1}=1 \mid \text{prior clicks}=k)$, as a function of the number of past clicks $k$. The weighted linear fit has a positive and statistically significant slope: users who have accumulated more clicks are more likely to click again on the next impression.\footnote{Final impressions (with no $t{+}1$) are excluded. These patterns are descriptive and do not imply causal effects.}

figure[figure omitted — 678 chars of source]

Taken together, these descriptive results reveal two dimensions of behavioral heterogeneity. First, Figure (ref) shows that CTR declines as exposure histories lengthen, indicating the need to account for user trajectories. Second, Figure (ref) shows persistence in click behavior: users with more prior clicks are more likely to click again, reflecting heterogeneity in underlying propensities and highlighting the predictive value of behavioral history.

Geographical Heterogeneity

We now turn to geographical patterns of responsiveness to examine whether user responses exhibit systematic spatial dependence. To do so, we group impressions by county and compute the average CTR within each spatial unit. Specifically, for a given county \( c \), we calculate \( \text{CTR}_c = \frac{1}{N_c} \sum_{i \in c} y_i, \) where \( y_i \in \{0, 1\} \) indicates whether impression \( i \) resulted in a click, and \( N_c \) is the number of impressions observed in county \( c \). Figure (ref) visualizes the resulting CTR values as a choropleth map.

figure[figure omitted — 170 chars of source]

The figure reveals pronounced spatial correlation in responsiveness: counties with high (low) average CTRs tend to be geographically proximate to other counties with similarly high (low) CTRs, which generate visible regional clusters rather than isolated pockets of responsiveness. This spatial structure indicates that user engagement is not independently distributed across space but instead varies smoothly across neighboring locations. Such correlation suggests that geographic context captures shared latent factors, such as local demographics or socioeconomic context, that influence users' likelihood of clicking on ads.

Train-Test Split Strategy

Finally, we split the data into training and test sets to evaluate out-of-sample predictive performance and policy effectiveness. To avoid information leakage, the split is performed at the user level rather than the impression level: each user’s full history is contained entirely within either the training or the test partition, allowing us to assess generalization to previously unseen users.

figure[figure omitted — 209 chars of source]

Figure (ref) illustrates the procedure. A user becomes eligible once they first appear in the platform’s logs, after which we track their impression history. Each user is then randomly assigned to either the training set (60%, 263,494 users; 6,202,022 impressions) or the test set (40%, 175,663 users; 4,134,681 impressions). Gray dots mark ineligible impressions that precede a user’s first appearance, while red and blue dots denote impressions assigned to the training and test sets, respectively. The resulting split maintains similar distributions of activity and click behavior across both partitions.

Problem Definition

Advertising platforms operate in environments where advertising decisions must be made under uncertainty about user preferences. Consider the case of an ad network deciding which ad to show to a particular user. The central goal of the platform is to maximize a reward measure, such as advertising revenue or user engagement. At the same time, the platform does not fully observe the user's type but observes a signal through observed features. As a result, the platform faces a decision-making problem under uncertainty, where the challenge lies in selecting the right action based on partial information about the user.

To analyze this decision problem, we model ad allocation as a choice among feasible actions in the presence of heterogeneous users and incomplete information. Users differ in latent preferences that determine the rewards generated by different ads, while the platform observes only an information set \(X\) that shapes its beliefs about these rewards. We now formalize the elements of this problem.

itemize• Environment. The platform must choose an action from a finite set $\mathcal{A}$, where each $a \in \mathcal{A}$ represents a feasible option (e.g., an ad that can be shown). Users differ in unobserved characteristics captured by a latent type $\omega \in \Omega$, with $\Omega$ denoting the space of possible user types. The reward of taking action $a$ for a user of type $\omega$ is represented by a function: \[ r: \mathcal{A} \times \Omega \to \mathbb{R}, \qquad r(a, \omega) \text{ (e.g., expected CTR)}. \] This reward function encodes the payoff of each action for each user type and serves as the platform’s objective. • Information set and user types. The platform does not observe the user's type $\omega$ directly. Instead, they observe an information set $X \in \mathcal{X}$ that provides partial knowledge about $\omega$. The information set $X$ may include behavioral variables, device attributes, geographical indicators, or other covariates. We assume the pair $(X, \omega)$ is drawn from a joint distribution $p$ over $\mathcal{X} \times \Omega$, which factorizes as \(p(x,\omega) = p(\omega)\,p(x\mid \omega)\), where $p(\omega) \in \Delta(\Omega)$ represents the population-level distribution of user types. Upon observing $X = x$, the platform forms a posterior belief over types: \[ q_x(\omega) := p(\omega \mid x) \in \Delta(\Omega). \] For notational clarity, we also define the random posterior $q_X := p(\cdot \mid X)$, a random variable that maps each realization $X = x$ to its corresponding posterior $q_x$. This posterior encodes the platform’s updated belief about the user's type after observing $x$. • Policies. A policy specifies how the platform selects actions based on the observed information set. Formally, a policy is a measurable function: \[ \pi: \mathcal{X} \to \mathcal{A}, \] which maps each realization $x \in \mathcal{X}$ to an action $\pi(x) \in \mathcal{A}$. We focus on deterministic policies because, under linear expected reward objectives, randomization offers no additional value (by an extreme-point argument; see kamenica2019bayesian, smith2023optimal)\footnote{When $\mathcal{A}$ is finite and the objective is linear, any mixed policy can be written as a convex combination of deterministic ones, and the maximum is achieved at an extreme point. Ties are resolved by an arbitrary but fixed rule.}. • Expected reward and optimal actions. Given a belief $q \in \Delta(\Omega)$ over user types, such as the posterior $q_x$ induced by observing $x$, the expected reward of choosing action $a \in \mathcal{A}$ is \[ \mathbb{E}_{\omega \sim q_x}[r(a, \omega)]. \] Following, among all possible policies $\pi \in \Pi$ that map observed information to actions, the optimal policy $\pi^*(x)$ is the one that maximizes expected reward for each observation: \begin{equation} \pi^*(x) \in \arg\max_{a \in \mathcal{A}} \ \mathbb{E}_{\omega \sim q_x}[r(a, \omega)]. \end{equation} • Decision value. Let $X$ denote the information set observed prior to action selection, and let $q_X = p(\cdot \mid X)$ be the induced posterior over user types. For any deterministic policy $\pi: \mathcal{X} \to \mathcal{A}$, its ex-ante value is \(V(\pi) = \mathbb{E}_{(X,\omega)\sim p}[r(\pi(X), \omega)]\). Therefore, the value of information from $X$ is the maximum achievable value over all such policies: \begin{equation} V_X = \sup_{\pi: \mathcal{X} \to \mathcal{A}} V(\pi) = \sup_{\pi: \mathcal{X} \to \mathcal{A}} \mathbb{E}_{X}\big[\mathbb{E}_{\omega \sim q_X}[r(\pi(X), \omega)]\big]. \end{equation} The optimal policy $\pi^*(x)$ defined in Equation (ref) attains this supremum by maximizing expected reward for each observation.

In this study, we focus on click outcomes (CTR) as the primary measure of reward. Clicks are central to digital advertising because they are closely tied to both platform revenue and user engagement under prevalent cost-per-click and click-weighted pricing mechanisms. They provide an immediate and observable measure of user response that reflects the quality of the platform’s ad allocation decision. Given this setup, we operationalize rewards using click outcomes.\footnote{We focus on clicks as the primary outcome because targeting value arises from increasing the relevance and clickability of ads. However, the framework readily extends to alternative reward measures, such as revenue per click.}

Let $Y \in \{0,1\}$ denote the binary indicator of user engagement (e.g., whether the user clicks on the displayed ad), and let $a \in \mathcal{A}$ denote the action chosen by the platform. Consistent with the model above, the realized outcome $Y$ depends on both the selected action $a$ and the user’s latent type $\omega \in \Omega$, which captures unobserved preferences and motivations. While $\omega$ is not directly observable, the platform may observe proxy information $X$ that is informative about it. We categorize available user-level information into two main types:

itemize• Behavioral Information ($X^B \in \mathcal{X}^B$) captures a user’s past interactions in the digital environment, which includes prior app usage patterns, ad exposure sequences, click history, and other engagement-based indicators that reflect individual preferences and interests over time. • Geographical Information ($X^G \in \mathcal{X}^G$) describes the user’s spatial context, encompassing attributes such as precise location (e.g., latitude/longitude) and broader administrative regions such as city/province that may correlate with demographic or regional characteristics.

We denote by $X^{GB}$ the information regime that combines both behavioral and geographical information. We also treat contextual information, denoted by $X^{\varnothing}$, as standard metadata attached to each ad impression, including device characteristics (e.g., brand and operating system) and network conditions (e.g., connectivity type). Like behavioral and geographical information, contextual variables reflect aspects of the user’s latent type $\omega \in \Omega$, but they differ in two important respects. First, they are typically available to platforms for every impression, making them a natural baseline information set. Second, consumers generally view them as substantially less privacy-sensitive than detailed behavioral or location-based data jerath2024consumers. Accordingly, we include contextual information ($X^{\varnothing}$) in all targeting regimes. The specific variables included in each regime $X \in \{X^{\varnothing}, X^{G}, X^{B}, X^{GB}\}$ are described in \S(ref).

While we formalize the value of information through its impact on decision outcomes in Equation (ref), a common alternative approach evaluates information solely through predictive performance. This approach focuses on estimating $\hat{y} = f(X,a)$ to approximate $\mathbb{P}(Y=1 \mid X,a)$ and assesses information quality using metrics such as AUC or reductions in conditional entropy $H(Y \mid X,a)$. However, predictive accuracy does not in general translate into better decisions or improved outcomes when targeting policies are deployed ascarza2018retention, rafieian2023ai.

Empirical Framework

The central objective of this framework is to assess the value of information by quantifying and comparing how behavioral data ($X^B$) and geographical data ($X^G$) contribute to improved decision-making. The value function in Equation (ref) provides the formal basis for this assessment by defining the maximum expected reward attainable under each information set. Implementing this framework, however, raises several challenges:

enumerate• Challenge 1: Information Value, Complementarity, and Substitutability: The first challenge is to formalize how to compare the value of distinct sources of information. When multiple data sources, such as behavioral (\(X^B\)) and geographical (\(X^G\)), are available, their value is inherently relative: each may add incremental benefit beyond the other or overlap in what it conveys. The marginal value of one source depends on whether the other is already observed, raising the possibility of complementarity or substitutability. Without precise definitions, we cannot separate the standalone contribution of each source from the value created by using them jointly. We therefore develop a formal framework that isolates individual value, incremental value conditional on the other source, and their combined effect. We present this framework in \S(ref). • Challenge 2: Reward Function Estimation for Decision Policies: The second challenge is empirical: estimating the reward function \( r(a,\omega) \). Our objective is to assess whether geographical information provides incremental decision value beyond behavioral information. Doing so requires reward estimation that fully captures the dynamic structure of user histories; otherwise, estimated gains from geographical information may reflect behavioral misspecification rather than informational value. Because behavioral information is inherently temporal, reward estimation must account for this dynamic structure. We address this challenge in \S(ref). • Challenge 3: Information Value Estimation through Policy Evaluation: The third challenge concerns estimating the value of information \(V_X\). By definition, \(V_X\) is the maximum expected reward achievable under a given information set, which requires comparing alternative decision policies. Such comparisons are inherently counterfactual, since each user is observed under only one realized action and outcomes under other actions are not observed. Estimating \(V_X\) therefore requires a model-free approach to evaluating the performance of counterfactual policies using logged data, even when some actions are rarely chosen. We address this challenge in \S(ref).

To address these challenges, we develop a unified framework that integrates economic theory, machine learning, and causal inference. The framework (i) formally defines information value to compare behavioral and geographical data and characterize complementarity or substitutability (\S(ref)), (ii) estimates reward functions that capture the dynamic structure of behavioral histories (\S(ref)), and (iii) evaluates information value through model-free counterfactual policy evaluation using logged data (\S(ref)). Together, these components quantify the decision value of behavioral and geographical information.

Substitutability and Complementarity of Information Sets

To address Challenge 1, we use formal definitions from the economic theory literature to compare the value of different information sets in a decision problem. In our context, we assess how behavioral data \(X^B\) and geographical data \(X^G\) contribute to guiding optimal actions. Intuitively, two pieces of information complement each other if the joint incrementality in information value goes beyond a simple sum of the incrementality of each alone.

We adopt the decision-theoretic framework of substitutability and complementarity introduced by borgers2013signals. Recall that \(X^B\) and \(X^G\) denote behavioral and geographical data, respectively. We define the combined information set as \(X^{GB} := (X^B, X^G) \in \mathcal{X}^B \times \mathcal{X}^G\), which represents joint access to both sources. As a baseline, we consider \(X^{\varnothing}\), corresponding to the absence of user-level information. In this case, decisions rely only on contextual information that informs the prior belief about user types, so the posterior reduces to the prior, \(q_{X^{\varnothing}} \equiv p(\omega)\). The associated value, \(V_{X^{\varnothing}}\), therefore represents decision-making under prior uncertainty without conditioning on individual-level signals. We define substitutability and complementarity by the marginal value of one information set conditional on access to the other, using the value function \(V_X\) (Equation (ref)).

definition[Substitutes] Behavioral information \(X^B\) is a substitute for geographical information \(X^G\) if, for every decision problem \footnote{Following borgers2013signals, “every” refers to the inequalities holding for all decision problems \((A,r)\). Our empirical analysis evaluates these same inequalities for a fixed problem \((A,r)\) (CTR objective and action set), yielding complement/substitute conclusions for this problem.} \((\mathcal A,r)\) , \[ V_{X^G} - V_{X^\varnothing} \;\ge\; V_{X^{GB}} - V_{X^B}. \]

This inequality compares the marginal value of geographical data when used alone versus when combined with behavioral data. The left-hand side, \( V_{X^G} - V_{X^\varnothing} \), reflects the value of geographical data on its own, while the right-hand side, \( V_{X^{GB}} - V_{X^B} \), reflects its incremental contribution when behavioral data are already available. If the inequality holds, behavioral data reduces the marginal usefulness of geographical data, making the two substitutes.

definition[Complements] Behavioral information \(X^B\) is a complement to geographical information \(X^G\) if, for every decision problem \((\mathcal{A},r)\), \[ V_{X^G} - V_{X^\varnothing} \;\le\; V_{X^{GB}} - V_{X^B}. \]

Here, the marginal value of geographical data is greater when combined with behavioral data than when used alone. The left-hand side, \(V_{X^{GB}} - V_{X^B}\), measures the incremental contribution of geographical data given that behavioral data are already available, while the right-hand side, \(V_{X^G} - V_{X^\varnothing}\), captures their standalone value. If the inequality holds, behavioral and geographical data enhance each other’s usefulness and are thus considered complements.

Machine Learning Framework for Reward Function Estimation

To address Challenge 2, we develop a machine learning framework for estimating expected reward when the reward function \( r(a, \omega) \) is unobserved. The ad platforms cannot observe latent types \( \omega \), nor directly measure the reward associated with each action. Instead, we re-express reward in terms of observable data: user characteristics \(X\), chosen actions \(A\), and realized engagement outcomes \(Y \in \{0,1\}\).

We specify reward through the binary click outcome \(Y\) and the reward value of engagement. Suppose that showing ad \( a \in \mathcal{A} \) to a user of type \( \omega \in \Omega \) yields reward only when the user clicks, and that the ad platform receives a known per-click reward \( v(a) \in \mathbb{R}_{\geq 0} \).\footnote{Since the targeting value of information arises from improved match quality and resulting changes in users’ CTR, we focus on this reward measure and set $v(a)=1$ for simplicity. All insights extend to settings in which $v(a)$ is heterogeneous, provided that CTR and per-click valuation are independent and separable.} In this case, the expected reward from taking action \( a \) for a user of type \( \omega \) is given by \[ r(a, \omega) := \mathbb{E}[Y \mid A = a, \omega] \cdot v(a) = \Pr(Y = 1 \mid A = a, \omega) \cdot v(a). \] This formulation captures both the probabilistic nature of engagement and the economic payoff per click, and it serves as the foundational definition of reward in our analysis. Although user types \( \omega \) are unobserved, conditioning on observed characteristics \( X = x \) induces a posterior belief \( q_x(\omega) := p(\omega \mid x) \) about the user’s latent type. Accordingly, we focus on estimating the function: \[ f_X(a, x) := \Pr(Y = 1 \mid X = x, A = a), \] which maps each action–characteristics pair to the probability of engagement. This function is the central object in our machine learning framework. Once \( f_X \) is estimated, the expected reward can be easily computed. This approach can be operationalized using flexible prediction models trained on observational data, provided that the model class is sufficiently expressive and the estimation procedure is properly regularized and calibrated.

To operationalize this framework, we estimate the conditional click function under several information regimes that differ in the richness of observable user data. Specifically, we consider four regimes \( \{X^\varnothing, X^G, X^B, X^{GB}\},\) which correspond to progressively richer access to contextual, geographical, and behavioral information. Each regime captures a distinct informational constraint faced by the advertising platform. The specific features included in each information set are described in Appendix \S(ref).

In the following, we describe learning algorithm selection \( f \) from a flexible class \( \mathcal{F} \) to approximate \( \Pr(Y = 1 \mid X, A) \) in \S(ref), and we specify the behavioral modeling architecture in \S(ref). The estimated function \( \hat{f}_X \) delivers action-specific rewards, induces policies \( \pi(x) \), and forms the basis for policy evaluation across informational regimes.

Learning Algorithm Selection Across Information Regimes

In regimes without behavioral histories (\(X^\varnothing\) and \(X^G\)), the information set is static across impressions. Because no new user-specific information accumulates over time, the estimation problem is inherently cross-sectional. We therefore employ flexible tree-based learners (XGBoost) to estimate the click probability as a function of contemporaneous covariates and actions. These models are well-suited to capturing nonlinearities and high-order interactions in static feature spaces without imposing strong parametric structure Rafieian2021.

By contrast, regimes that include behavioral information (\(X^B\) and \(X^{GB}\)) generate sequential user histories whose informational content expands over time. In these settings, the relevant state variable is no longer a fixed vector but an evolving sequence of past exposures and engagement outcomes. To accommodate this structure, we develop a sequence model based on the LSTM network with an attention mechanism, which aggregates information over time and forms latent summaries of user behavior quadrana2018sequence. These models map behavioral histories and current actions into predicted engagement probabilities, allowing the estimator to adapt dynamically as additional impressions are observed.

This distinction between static and sequential regimes is central to our analysis. As behavioral information accumulates over impressions, the information set \(X^B_{jt}\) becomes increasingly informative about user behavior. This richer information allows the advertiser to condition decisions on a more precise representation of past engagement, improving estimation of the engagement function \(f(X^B_{jt}, a)\). As a result, targeting decisions become more accurate, and the expected reward under the induced policy increases as behavioral histories grow. We examine this empirically in \S(ref).

Sequence-Based Behavioral Modeling with LSTM and Attention

We now describe the sequence model used for regimes that include behavioral histories in this part. We design an LSTM architecture equipped with a causal multi-head attention mechanism. The LSTM captures short- and long-range temporal dependencies in user interactions, while attention highlights the most informative parts of a user’s exposure–click history. Together, these components allow the model to represent both persistent behavioral trends and short-term recency effects, which are central to engagement modeling quadrana2018sequence.

We train the model on user interaction histories using a sliding window of 150 impressions. At each time step, the input combines categorical and numerical features. We map high-cardinality categorical variables (e.g., device model, app ID, network identifiers) into dense vectors with embedding layers guo2016entity and concatenate them with continuous covariates. To encode order, we add absolute positional embeddings vaswani2017attention. We also process the log-transformed inter-arrival time through a time-gap projection, which helps the model detect irregular timing patterns that often indicate behavioral shifts.

figure[figure omitted — 202 chars of source]

We feed the enriched sequence into a four-layer LSTM with hidden size 512. On top of the LSTM outputs, we leverage causal multi-head self-attention with four heads. The attention module uses a future-masked matrix to block access to unseen impressions, which preserves temporal causality and prevents information leakage vaswani2017attention. We apply layer normalization before and after the attention block to stabilize training. Next, we add a gated projection head that runs parallel sigmoid and tanh transformations, combines them elementwise, and applies dropout for regularization. A fully connected output layer finally maps the hidden states to click probabilities at each step.

Figure (ref) illustrates the designed sequence model architecture and its key components. The design directly leverages sequential order, irregular timing, and varying relevance of past impressions. As users generate longer histories, the LSTM–attention framework provides richer information about preferences and latent types. We provide additional implementation details in Web Appendix \S(ref).

Estimating the Value of Information via Inverse Propensity Scoring

Challenge 3 highlights the core difficulty in estimating the value of information $V_X$(Equation (ref)): \[ V_X = \sup_{\pi: \mathcal{X} \to \mathcal{A}} \mathbb{E}_X\left[\mathbb{E}_{\omega \sim q_X}[r(\pi(X), \omega)]\right], \] The data only reveal outcomes under the action chosen by the platform’s logging policy, while $V_X$ requires utilities aggregated over all possible actions. Each impression, therefore, provides information on one action, but evaluating a policy involves counterfactual outcomes for actions that were not taken.

We address this problem with Inverse Propensity Scoring (IPS), a standard method in causal inference for adjusting observational data hirano2003efficient, rafieian2023ai. IPS identifies the value of a fixed target policy by reweighting observed outcomes according to how likely the logging policy was to select the same action as the target policy. By doing so, we approximate the reward that would have been realized under alternative policies, even though those policies were never deployed in practice.

Formally, for each information set $X$, we define a target policy $\pi^{f_X}$ that deterministically selects the action maximizing the estimated expected reward based on our predictive model: \[ \pi^{f_X}(a \mid x) \;=\; \mathbb{I}\!\left\{a \;=\; \arg\max_{a' \in \mathcal{A}} \hat{f}_X(a, x)\right\}, \] where $\hat{f}_X(a, x)$ is the estimated click probability given information set $x$ and candidate action $a$. Thus, $\pi^{f_X}$ prescribes, for each instance, the action expected to maximize reward under the model. The value of this policy, consistent with our decision-theoretic objective, is

\[ V(\pi^{f_X}) \;=\; \mathbb{E}_X\!\left[\mathbb{E}_{\omega\sim q_X}\big[r(\pi^{f_X}(X),\omega)\big]\right]. \] However, only outcomes corresponding to the logging policy’s actions are observed in the data. To estimate the value of the induced policy $V(\pi^{f_X})$, we use the IPS estimator, which reweights the observed reward outcomes for each user-impression pair by the inverse of the probability that the logging policy would have selected the action recommended by the target policy: \[ \hat V(\pi^{f_X}) = \frac{1}{N}\sum_{i=1}^N \frac{\mathbb{I}\!\big[A_i=\pi^{f_X}(X_i)\big]}{\pi^{\mathcal D}(A_i\mid X_i)} \, v(A_i)\, Y_i, \] where $A_i$ is the action taken in the historical data for impression $i$, $\pi^{\mathcal{D}}(A_i \mid X_i)$ is the probability that the logging policy selected $A_i$ given information $X_i$, and $Y_i$ is the observed binary engagement outcome. Each term, therefore, reweights the realized reward to reflect how the target policy would allocate actions. For the IPS estimator to recover the true value $V_X$, three key conditions must hold regarding the data-generating process and the structure of the logging policy. We state these assumptions formally below.

assumption[Overlap] For all $x \in \mathcal{X}$ and $a \in \mathcal{A}$, if $\pi^{f_X}(a \mid x) > 0$, then $\pi^{\mathcal{D}}(a \mid x) > 0$. In other words, the logging policy must assign positive probability to every action that the target policy might select.
assumption[Unconfoundedness] Potential outcomes are independent of the action actually taken, conditional on observed covariates $X$; that is, $Y(a) \perp A \mid X$ for all $a \in \mathcal{A}$. This ensures that all confounding factors are captured in $X$.
assumption[Policy optimality under $X$] If $\hat f_X$ consistently estimates $\Pr(Y=1\!\mid\!A=a,X=x)$ and the maximizer of $v(a)\Pr(Y=1\!\mid\!A=a,X=x)$ is unique almost everywhere, then $\pi^{f_X}(x)=\arg\max_a v(a)\hat f_X(a,x)$ is optimal under $X$ almost everywhere.
proposition[Standard IPS identification and recovery of $V_X$] Under Assumptions (ref)--(ref), and when $\pi^{f_X}$ is evaluated on an independent sample, \[ \mathbb{E}\!\big[\hat V(\pi^{f_X})\big] \;=\; V_X, \qquad \hat V(\pi^{f_X}) \xrightarrow{p} V_X. \] That is, the IPS estimator is unbiased and consistent for $V_X$ under the stated conditions. \\ \begin{proof} A formal proof is provided in Web Appendix \S(ref). \end{proof}

The IPS approach addresses Challenge 3 by enabling estimation of the ex-ante value of any information set \( X \) through the policy it induces, even when counterfactual outcomes are unobserved. In the following, we first discuss the conditions and diagnostics required to validate propensity score estimation in \S(ref), and then describe the empirical procedure used to estimate propensity scores in \S(ref).

Estimating and Validating Propensity Scores

To implement the IPS estimator, we require estimates of the logging policy probabilities \( \pi^{\mathcal{D}}(a \mid x) \), which govern how actions are assigned in the historical data. First, we describe the ad allocation mechanism that generates these probabilities, and then examine empirical evidence supporting the IPS identification assumptions.

\paragraph{Quasi-Proportional Allocation Mechanism.} In our setting, impressions are allocated via a quasi proportional auction that induces randomized exposure across eligible ads as a function of observed bids, quality scores, and eligibility constraints mirrokni2010quasi. For each impression \( i \), let \( \mathcal{A}_i \subseteq \mathcal{A} \) denote the set of ads eligible to participate in the auction. Each ad \( a \in \mathcal{A}_i \) submits a bid \( b_a \) and has a platform-assigned quality score \( q_a \), both of which are fixed and observed during our sample period. The probability that ad \( a \) wins the auction and is shown in impression \( i \), conditional on covariates \( x_i \), is given by: \[ \pi^{\mathcal{D}}(a \mid x_i) = \frac{b_a q_a}{\sum_{j \in \mathcal{A}_i} b_j q_j}. \] This allocation rule induces a probabilistic assignment over eligible ads, with probabilities that are fully determined by observed features. In contrast to deterministic formats (e.g., second-price auctions) where only the top-ranked ad is observed, the quasi-proportional mechanism introduces randomization across all eligible ads. This variation is central to our empirical framework, as it enables estimation of counterfactual outcomes and supports off-policy evaluation.

remark[Overlap] Any ad participating in the auction for impression \( i \) (i.e., \( \forall a \in \mathcal{A}_i \)) has a nonzero propensity of being shown in impression \( i \).

This follows directly from the quasi-proportional rule: every ad with a positive \( b_a q_a \) has a strictly positive probability of being selected. Hence, the overlap assumption (ref) is satisfied by construction for all participating ads in each auction.

remark[Unconfoundedness] For any impression \( i \), ad allocation is independent of the set of potential outcomes for participating ads \( (a \in \mathcal{A}_i) \), after controlling for the observed covariates. Thus, \[ \{Y_i(a)\}_{a \in \mathcal{A}_i} \perp A_i \mid x_i. \]

Therefore, the unconfoundedness assumption (ref) is satisfied by the transparent structure of the auction: all inputs that determine allocation, namely, bids \( b_a \), quality scores \( q_a \), and eligibility constraints, are fully observed and fixed over the sample period. For each impression \( i \), we observe the complete covariate vector \( x_i \). Advertiser bids are not dynamically adjusted, and the platform does not personalize quality scores across users. As a result, the assignment rule is fully determined conditional on \( x_i \).

This structure implies that for any impression \( i \), we can compute not only the probability \( \pi^{\mathcal{D}}(A_i \mid x_i) \) of the ad that was actually shown, but also the probabilities of all counterfactual ads that were eligible in the same auction. That is, even if a particular ad \( a \) was not displayed in impression \( i \), as long as it was eligible (i.e., \( a \in \mathcal{A}_i \)), we can recover the counterfactual allocation probability \( \pi^{\mathcal{D}}(a \mid x_i) \).

\paragraph{Eligibility Filtering.} Although the quasi-proportional rule defines probabilities \( \pi^{\mathcal{D}}(a \mid x_i) \) for all ads in the auction, not all ads in the global set \( \mathcal{A} \) are eligible in every impression. To ensure valid counterfactual estimation, we restrict attention to ads with nonzero probability of participation in each impression. Two factors determine eligibility: (i) Contextual Targeting: ads may be restricted to specific provinces, times, or app categories, and are excluded when their targeting criteria are not met; (ii) Campaign availability: some ads may be inactive due to budget exhaustion or campaign timing. In practice, this is rare among top ads, as we select the top 10 ads for our analysis.

We construct an eligibility matrix \( E \in \{0,1\}^{N \times A} \), where \( e_{i,a} = 1 \) indicates that ad \( a \) was eligible to compete in impression \( i \), based on observed targeting and availability constraints. In practice, this requires that the impression’s metadata match the ad’s targeting filters on province, hour-of-day, and app. To avoid misclassification, we drop any impressions with missing targeting variables and restrict attention to a filtered sample where eligibility can be verified.

While this filtering step identifies the support of the logging policy, it does not tell us how likely each eligible ad is to be selected. For unbiased off-policy evaluation, we must account for the non-random assignment probabilities across eligible ads. Next, we describe how we estimate these propensities and assess the validity of the unconfoundedness assumption via covariate balance.

Propensity Score Estimation and Covariate Balance

To correct for unequal selection probabilities inherent in the auction, we estimate the propensity scores \(\pi^{\mathcal{D}}(a \mid x_i)\), defined as the probability that the logging policy assigns impression \(i\) to ad \(a\) given the observed features \(x_i\). These propensities quantify the exposure pattern generated by the platform’s allocation mechanism and form the basis for IPS in our policy evaluation.

Although the quasi-proportional rule maps bids and quality scores into theoretical probabilities, we adopt a data-driven estimation strategy to capture the realized assignment process in practice. This approach flexibly accommodates deviations from the theoretical rule, nonlinear effects, and high-order interactions among features. We use XGBoost to estimate \(\pi^{\mathcal{D}}(a \mid x_i)\), motivated by its empirical performance in high-dimensional classification problems rafieian2023ai. To avoid overfitting and ensure out-of-sample validity, we implement a 5-fold cross-fitting procedure, so that each estimated propensity score is computed on a model trained without the corresponding observation.

While observed inputs fully determine the allocation rule, we validate our identifying assumption by testing whether inverse-propensity weighting balances the distribution of observed covariates across treatment. The variables examined include province, app context, time of day, device brand, network type, and mobile service provider, features that advertisers can directly target. For each covariate and treatment group, we compute the standardized mean difference (SMD) before and after applying the estimated propensity weights, considering absolute values below 0.2 as indicative of acceptable balance mccaffrey2013tutorial. The diagnostics show substantial improvements in covariate alignment across ads after weighting, providing empirical support for our research design. Details of propensity score estimation and balance statistics are reported in Web Appendix \S(ref).

Empirical Results

We now present the empirical results derived from our proposed framework. \S(ref) reports the predictive performance of machine learning models trained on different information sets. We then turn to the core question of how behavioral and geographical information contribute to decision quality, examining it at two levels of analysis. First, at the aggregate level (\S(ref)), we quantify the overall value of each information set and test whether the two act as substitutes or complements in improving targeting performance. Second, we extend the analysis to the user level (\S(ref)), examining how the value and interaction of these information types change with the amount of behavioral history observed by the platform.

Predictive Performance Across Information Sets

Predicting user engagement is inherently challenging due to extreme class imbalance: fewer than 2% of impressions generate a click, making accuracy an uninformative metric. We therefore evaluate predictive performance using log loss, which assesses the calibration of predicted click probabilities and aligns with the models’ training objective; relative information gain (RIG), which measures the proportional improvement in log loss relative to a baseline model that predicts the average click-through rate; and the area under the ROC curve (AUC), which captures the model’s ability to rank clicked above non-clicked impressions in a threshold-independent and imbalance-robust manner.

We evaluate four models trained under different informational regimes: (i) the contextual-only baseline (\(X^{\varnothing}\)), (ii) the geographical regime (\(X^G\)), (iii) the behavioral regime (\(X^B\)), and (iv) the full information set (\(X^{GB}\)). The performance of each model on the held-out test set is summarized in Table (ref).

table[table omitted — 856 chars of source]

The results reveal substantial variation in predictive performance across information structures. The contextual-only baseline performs weakest, with an RIG of 18.90%. Adding geographical features yields only a modest improvement, raising RIG to 19.22% and AUC to 0.722. In contrast, incorporating behavioral information leads to a large performance gain: the behavioral model achieves an RIG of 84.18% and an AUC of 0.809, indicating both substantially richer information and a superior ability to rank click outcomes. Adding geographical information to the behavioral model produces only marginal additional improvements, suggesting that behavioral features account for most of the predictive variation.

Taken together, these results demonstrate the dominant predictive value of behavioral information in estimating click likelihood, which aligns with prior findings in the advertising Rafieian2021. While comparable AUC levels can be achieved using XGBoost-based behavioral models, the LSTM architecture delivers substantially higher RIG, indicating improved probabilistic calibration through temporal modeling. These performance levels provide empirical support for Assumption (ref). Additional comparisons between LSTM and XGBoost models are reported in Appendix \S(ref).

That said, it is important to emphasize again that predictive performance does not always translate into decision value. A model may be well-calibrated or rank instances correctly, yet offer limited benefit when used to guide actions under realistic constraints. We therefore turn next to assessing the actual targeting value generated by each model in terms of the value function.

Behavioral vs.\ Geographical Information: Aggregate Level

Having established model performance in \S(ref), we now assess how behavioral and geographical information improve decision quality. As described in \S(ref), we evaluate each targeting policy \( \pi^{(f_X)} \) using IPS. We proceed in two steps. First, in \S(ref), we quantify the aggregate value of each information set by comparing the expected value of policies based on different data sources. Then, in \S(ref), we test whether behavioral and geographical information act as substitutes or complements in shaping decision quality.

Aggregate Value of Behavioral and Geographical Information

We consider a set of deterministic greedy policies indexed by the information set \( X \in \{X^{\varnothing}, X^G, X^B, X^{GB}\} \) available at the time of decision. For each impression \( i \), the policy selects, among eligible actions \( \mathcal{A}_i \) (see \S(ref)), the ad with the highest predicted click probability, restricting attention to actions with positive estimated logging propensity \( \hat{\pi}^{\mathcal D}(a \mid X_i) > 0 \). This ensures that all selected actions lie within the support of the logging policy and satisfy the overlap condition in Assumption (ref).

Each policy represents a counterfactual targeting scenario in which the advertiser optimizes using only the corresponding information set \( X \), and its value \( \hat V_X \) is estimated via IPS. To assess estimator uncertainty, we report 95% confidence intervals based on cluster-robust standard errors (clustered at the user level) and the effective sample size (ESS), which captures variance inflation due to skewed importance weights kallus2019stable.

The results are reported in Table (ref). Across all policies, we find statistically significant improvements over the historical logging baseline, with tight confidence intervals and high ESS. In particular, mccaffrey2013tutorial advocates trimming weights until \(\mathrm{ESS} \geq 0.10N\), a threshold all policies in our setting easily exceed. This suggests that the estimated policy values are stable and inference is well-powered.

table[table omitted — 1,180 chars of source]

Starting from the contextual information policy \( X^\varnothing \), we observe a 20.4% improvement over the baseline CTR. Although the platform lacks access to user-level data, this policy leverages contextual variation and selects the ad with the highest average CTR. This result highlights the value of exploiting aggregate performance differences even in the absence of personalized information.

Access to richer information sets produces further gains in policy value. Geographical information (\( X^G \)) and behavioral information (\( X^B \)) both generate substantial improvements relative to the contextual policy, increasing lift to 28.7% and 30.3%, respectively. Combining both information sources (\( X^{GB} \)) yields the highest policy value, a 41.5% improvement over the logging baseline. Notably, unlike the predictive results in \S(ref), where behavioral information explained most of the performance gain, geographical information contributes comparably to decision value. This pattern underscores that improvements in predictive accuracy do not necessarily translate into proportional improvements in decision quality. Additional comparisons across learning algorithms are reported in Web Appendix \S(ref).

Complement or Substitute? Aggregate Level

We now investigate whether behavioral and geographical data act as substitutes or complements in informing ad targeting decisions, starting with the aggregate level. Formally, let: \[ \Delta \;=\; \big(\hat{V}_{X^{GB}} - \hat{V}_{X^B}\big) - \big(\hat{V}_{X^G} - \hat{V}_{X^\varnothing}\big), \] where $\hat{V}_{X}$ denotes the IPS-based value estimate for policy $X$. The first term in parentheses measures the incremental value of adding geographical information when behavioral data are already available, while the second term measures the incremental value of adding geographical information when no user-level data are observed.

If $\Delta > 0$, behavioral and geographical information are complements (Definition (ref)), meaning the value of combining them exceeds the sum of their stand-alone gains. If $\Delta < 0$, they are substitutes (Definition (ref)), meaning the combined gain is less than additive and the two information sets overlap in the information they provide. We estimate $\Delta$ using per-impression differences and conduct a one-sample $t$-test on the mean, clustering standard errors at the user level to account for correlation within users. We report the two-sided test, as one-sided results cannot be significant when the two-sided test fails to reject the null.

table[table omitted — 702 chars of source]

As shown in Table (ref), the aggregate interaction estimate is close to zero and statistically insignificant. This implies that, on average, combining behavioral and geographical data produces nearly additive gains, with no systematic evidence of complementarity or substitutability at the aggregate level. While each information set captures distinct user heterogeneity, their joint effect does not amplify or diminish targeting value in the aggregate. Given this null aggregate finding, we next examine whether the nature of the interaction varies with user impression depth to assess whether heterogeneity in complementarity or substitutability emerges across users.

Behavioral vs.\ Geographical Information: Heterogeneity by User Exposure

The null result at the aggregate level motivates a closer examination of how the value and interaction of behavioral and geographical information vary across users. We focus on heterogeneity by impression depth, the number of impressions observed from each user, which proxies how much the platform has learned about each user. We proceed in two steps. First, in \S(ref), we assess how the decision value of each information set changes with user exposure. Then, in \S(ref), we test whether the relationship between behavioral and geographical information shifts from substitutive to complementary as more behavioral data accumulate.

Heterogeneous Value of Behavioral and Geographical Information by User Exposure

Building on the aggregate results presented in \S (ref), we now explore how the value of information varies with the user's impression history. Specifically, we investigate whether the gain from targeting differs depending on how many impressions a user has previously seen. This analysis is motivated by the discussion in \S(ref), which posits that the estimation of reward function \( f(X^B_t, a) \) becomes more accurate as more behavioral data is observed. If true, the marginal value of behavioral or combined data may evolve with impression depth.

To operationalize this, we first sort all impressions for each user \( j \) by their timestamp \( t \) and assign a depth index accordingly. We then group observations into bins such that each bin contains the same number of impressions (i.e., quantile-based binning). Within each bin, we compute the absolute CTR levels for each targeting policy using the IPS-estimated click probabilities, alongside the empirical baseline CTR. These per-bin averages are then plotted against impressions have seen by user to visualize how click-through performance evolves as users receive additional exposures. Figure (ref) and Figure (ref) present these results for the main targeting comparisons.

figure[figure omitted — 786 chars of source]

In the left panel of Figure (ref), the \(X^G\) (Geographical) policy consistently delivers higher CTR than the \(X^\varnothing\) (Contextual) policy across all impression depths. This persistent gap indicates that location-based segmentation provides a stable improvement over contextual cues alone. Although both policies rely on static features, geographical information captures cross-regional heterogeneity that enhances targeting precision even in the absence of behavioral data.

In the right panel of Figure (ref), both the \(X^B\) (Behavioral) and \(X^{GB}\) (Combined) policies achieve substantially higher CTR than the baseline, highlighting the value of behavioral information for personalization. The \(X^{GB}\) policy consistently outperforms \(X^B\), though the gap narrows as impression depth increases. This pattern implies that while geographical data initially enhances personalization in early exposures, its marginal contribution diminishes once rich behavioral histories accumulate.

Complement or Substitute? Heterogeneity by User Exposure

figure[figure omitted — 443 chars of source]

Building on the aggregate results in Table (ref), we now examine how the relationship between $X^B$ and $X^G$ varies with the number of impressions a user has seen. We group impressions into bins of equal size, each containing the same number of observations, defined as the number of prior ads shown to the same user. For each bin, we compute two policy-value differences: (i) the value difference between the $X^{GB}$ and $X^B$ policies, and (ii) the value difference between the $X^{G}$ and $X^{\varnothing}$ policies.

Figure (ref) summarizes these patterns. The top panel plots both value differences across impression depth with 95% confidence intervals, while the bottom panel reports their difference-in-differences, \( \bar{\Delta} = (X^{GB} - X^B) - (X^{G} - X^{\varnothing}),\) where positive values indicate complementarity (joint value exceeds additivity) and negative values indicate substitutability (information sets overlap in value). Shaded regions denote ranges with statistically significant effects.

The figure reveals three stages. In the minimal behavioral history stage (approximately 1--2 impressions), \(X^B\) contains little behavioral information. As a result, the stand-alone contribution of geographical data \(X^{G} - X^{\varnothing}\) is at its highest. In this range, adding geographical data to behavioral data delivers nearly the same gain as geography alone, so \(X^{GB} - X^{B} \approx X^{G} - X^{\varnothing}\). Consequently, \(X^{GB}\) yields little additional value relative to \(X^{G}\).

In the sparse behavioral history stage (up to roughly 25 impressions), geographical information continues to play an important role as \(X^B\) begins to capture meaningful signals that, when combined with \(X^G\), create synergy between the two information sources. In this range, geographical and behavioral information act as complements: the value difference \(X^{GB} - X^B\) exceeds \(X^{G} - X^{\varnothing}\), which indicates that adding \(X^G\) to \(X^B\) delivers incremental targeting value beyond their separate contributions.

Finally, in the rich behavioral history stage (from the low 20s and beyond), the marginal contribution of geographical data declines as accumulated behavioral information becomes substantially more informative. In this regime, the marginal contribution of adding \(X^G\) to \(X^B\) turns negative, indicating substitutability, as \(X^B\) already captures much of the variation that \(X^G\) would otherwise provide. For the full set of statistical test results underlying this analysis, we refer readers to Web Appendix \S(ref).

Mechanisms Underlying the Role of Geographical Data

We now turn to the mechanisms that explain why geographical data ($X^G$) contributes to targeting outcomes in relation to behavioral data ($X^B$). Our unified framework has shown how the two information sets act as complements or substitutes in the decision value they generate. What remains unresolved is why geographical data improves performance in certain cases: does $X^G$ capture an independent channel of information about user responsiveness, or does it primarily serve as a proxy for preference patterns that behavioral data eventually reveal?

To address this, we decompose spatial correlation in ad responsiveness into two distinct sources. The first is influence, which has a causal interpretation: users respond similarly because they are in the same location, which enables peer interaction and local information transmission. The second is confounding: users respond similarly not because of the same location itself, but because location is confounded with other factors, such as similar characteristics, income, demographics, or cultural norms, that independently shape preferences. This distinction mirrors the social network literature, which distinguishes social influence from latent homophily as alternative explanations for correlated behavior among connected individuals anagnostopoulos2008influence, ma2015latent. We present the two sources of existing spatial correlation as follows:

enumerate• Spatial influence. Under the causal interpretation, users may affect the decisions of others around them. For example, a user who sees an in-app advertisement for a product or promotion may mention it to friends or coworkers nearby, increasing their likelihood of engaging with the same ad when they encounter it. In this case, correlated responses arise from information spreading through local interactions rather than from shared preferences. Detecting such effects requires observing the spatial network that connects individuals. Geographical data ($X^G$) provides this information by revealing which users are located near one another and therefore more likely to interact. Behavioral data ($X^B$), which records only each user's actions, cannot capture these local interaction patterns. Consequently, if spatial influence exists, it can only be detected through geographical data that allows the underlying spatial network to be constructed. • Spatial confounding. Under the confounding interpretation, users respond similarly not because they influence one another, but because location is correlated with shared characteristics that shape preferences. For example, residents of a wealthy neighborhood may be more likely to click on luxury-product ads because they have similar income levels and consumption preferences. Geographical data ($X^G$) captures this pattern immediately, since location acts as a proxy for the bundle of characteristics associated with that area. However, these preferences are gradually revealed through individual behavioral data ($X^B$). A user who repeatedly engages with premium brands or high-end retailers reveals the same underlying preference that geography initially proxies. As behavioral histories accumulate, models can infer these preferences directly from past actions, making geographical data primarily an early but coarse proxy whose value declines as behavioral data becomes richer.

Confounding in location intuitively means that the location selection by individuals is a function of underlying characteristics that are also correlated with the outcome of interest. To formalize this, recall $\omega \in \Omega$ that denote the true latent type of a user defined in \S(ref). The true click propensity is therefore a function of this latent type, which we denote by $g(\omega)$. Since $\omega$ is unobserved, the platform relies on observable behavioral data $X^B$ to estimate this function. Let $f(X^B)$ denote the platform's estimate of $g(\omega)$ based on the available behavioral history. We can therefore write click propensity as

equation[equation omitted — 190 chars of source]

where $\varepsilon$ captures idiosyncratic variation in responses. In Equation ((ref)), spatial influence is captured by $\theta_I$, while spatial confounding appears in the estimation error $\theta_C = g(\omega)-f(X^B)$. Geographical data $X^G$ is informative when spatial correlation remains in these terms after conditioning on behavioral data $X^B$. Under spatial influence, users affect the responses of nearby users through local interactions. These effects cannot be disentangled from individual preferences, regardless of how rich $X^B$ becomes. If influence exists, it creates spatial correlation in $\theta_I$ that $X^G$ can capture. Under spatial confounding, location is correlated with latent characteristics in $\omega$ that affect ad responsiveness. When behavioral histories are limited, $f(X^B)$ does not fully capture these characteristics, leaving spatial variation in $\theta_C=g(\omega)-f(X^B)$ that $X^G$ can explain. As behavioral histories accumulate and $f(X^B)$ better approximates $g(\omega)$, this component shrinks. The key empirical question is whether the predictive value of $X^G$ disappears as behavioral histories become richer or persists after conditioning on $X^B$.

Empirical Strategy: Residualized Spatial Autocorrelation Test

To evaluate this decomposition, we test whether spatial correlation remains after conditioning on behavioral data $X^B$. In Equation ((ref)), spatial confounding enters through $\theta_C = g(\omega)-f(X^B)$ and should diminish as behavioral histories allow $f(X^B)$ to better approximate $g(\omega)$. Spatial influence, captured by $\theta_I$, instead arises from local interactions and therefore persists regardless of how rich $X^B$ becomes. Our empirical strategy therefore tests whether residual spatial correlation disappears once behavioral information is incorporated. If it does, geographical data primarily reflects spatial confounding already absorbed by $X^B$; if it persists, it indicates spatial influence. We implement this idea using a residualized spatial autocorrelation (RSA) test.

Our empirical strategy proceeds by aggregating outcomes at the regional level. Let $Y_c$ denote the total number of clicks and $I_c$ the total number of impressions in region $c$. Each impression can be viewed as a Bernoulli trial with region-specific click probability $p_c$, but because click events are rare, the Binomial distribution is well approximated by a Poisson model: \[ Y_c \mid p_c \sim \text{Poisson}(I_c \, p_c). \] This specification is standard in spatial count-data analysis. The offset term $\log(I_c)$ adjusts for heterogeneous exposure across regions, while the residuals from the fitted model provide a natural basis for testing whether unexplained variation in click-through rates is spatially correlated.

list{$\bullet$} { {0pt} {1pt} {1pt} {1pt} {1.5em} {1em} {0.5em} } • Step 1: Baseline spatial structure in raw CTR: As a starting point, we examine whether CTRs exhibit geographical clustering in the absence of any behavioral controls. We estimate a Poisson count model that includes only an exposure offset and define the residual: \[ Y_c \sim \text{Poisson}(\mu_c), \quad \log \mu_c = \alpha + \log(I_c), \quad {\varepsilon^{(0)}_c = \log(Y_c/I_c) - \alpha}, \] Here, $\log(I_c)$ ensures that $\mu_c / I_c$ corresponds to the expected CTR in region $c$, while $\varepsilon^{(0)}_c$ measures deviations from the global mean CTR after adjusting for exposure volume. Evidence of spatial autocorrelation in $\varepsilon^{(0)}_c$ would indicate that raw CTRs contain location-based structure, consistent with spatial confounding or influence. • Step 2: Residual spatial structure after controlling for behavior: Next, we ask whether this spatial structure remains once behavioral data are taken into account. To do so, we augment the model with the fitted behavioral component $f(X^B_c)$ and compute the adjusted residual: \[ Y_c \sim \text{Poisson}(\mu_c), \quad \log \mu_c = \alpha + f(X^B_c) + \log(I_c), \quad {\varepsilon^{(B)}_c = \log(Y_c/I_c) - \alpha - f(X^B_c)}, \] This residual captures variation in CTR unexplained by either exposure or behavioral features. If $f(X^B_c)$ successfully absorbs location-linked patterns, spatial dependence in $\varepsilon^{(B)}_c$ should be weaker than in $\varepsilon^{(0)}_c$. • Step 3: Testing for spatial dependence: Finally, we test whether $\varepsilon^{(0)}_c$ and $\varepsilon^{(B)}_c$ exhibit spatial autocorrelation. We apply two standard measures: Moran’s $I$ moran1950notes, which captures global correlation across regions, and Geary’s $C$ geary1954contiguity, which is more sensitive to local clustering. For a generic set of residuals $\varepsilon_c$ indexed by region $c$, these are defined as: \[ I = \frac{N}{\sum_{i} \sum_{j} w_{ij}} \cdot \frac{\sum_{i} \sum_{j} w_{ij} (\varepsilon_i - \bar{\varepsilon})(\varepsilon_j - \bar{\varepsilon})} {\sum_{i} (\varepsilon_i - \bar{\varepsilon})^2}, \quad C = \frac{(N - 1) \sum_{i} \sum_{j} w_{ij} (\varepsilon_i - \varepsilon_j)^2} {2 \sum_{i} (\varepsilon_i - \bar{\varepsilon})^2}, \] where $N$ is the number of regions, $w_{ij}$ denotes the $(i,j)$ element of the spatial weight matrix $W$, and $\bar{\varepsilon}$ is the mean of the residuals. Significant positive values of Moran’s $I$ (or values of Geary’s $C$ below one) indicate that geographically proximate regions have similar residuals. By comparing the statistics for $\varepsilon^{(0)}_c$ and $\varepsilon^{(B)}_c$, we can assess whether behavioral features reduce spatial dependence, thereby clarifying whether geographical data provides independent information beyond what behavior explains.

Results: Spatial Correlation in CTRs

We present two sets of results. First, in \S(ref), we examine spatial correlation in baseline versus behavior-adjusted CTR residuals. Second, in \S(ref), we compare spatial dependence under sparse versus rich behavioral histories. Together, these analyses show how the role of geographical data changes once behavioral information is taken into account.

Residual Spatial Dependence in Baseline and Behavior-Adjusted Models

We begin by applying the RSA framework introduced in \S(ref). Specifically, we compare spatial correlation in the baseline residuals \(\varepsilon^{(0)}_c\), obtained without behavioral controls, to the behavior-adjusted residuals \(\varepsilon^{(B)}_c\). This analysis is conducted at both the county and city levels to assess how much of the observed spatial dependence can be explained by behavioral data.

figure[figure omitted — 397 chars of source]

Figure (ref) provides an intuitive visualization for the county-level case. The left panel displays the baseline residuals $\varepsilon^{(0)}_c$ from the model without behavioral controls. Several contiguous regions show clusters of similarly high or low residual values, indicating clear spatial autocorrelation. Such clustering is consistent with the presence of spatial homophily and/or influence. In contrast, the right panel shows the behavior-adjusted residuals $\varepsilon^{(B)}_c$. Incorporating $f(X^B_c)$ visibly reduces the strength and extent of spatial clusters, particularly in the southwest, although notable pockets remain. This pattern suggests that behavioral data absorb a substantial share of the location-linked variation, but not all of it.

To quantify the patterns observed in Figure (ref), Table (ref) reports Moran’s $I$ and Geary’s $C$ at both county and city levels. At the county level, the baseline residuals $\varepsilon^{(0)}_c$ exhibit strong and significant spatial dependence: regions with unusually high (or low) CTR residuals tend to be geographically proximate. After controlling for behavioral information, both statistics decline markedly, indicating that much of this clustering is explained by differences in user behavior. However, the remaining positive spatial correlation implies that geographical data continues to capture additional variation not yet absorbed by behavior.

At the city level, the attenuation is sharper. Baseline residuals still display detectable clustering, but once behavioral features are included, both Moran’s $I$ and Geary’s $C$ fall to levels that are statistically indistinguishable from zero. This result suggests that at finer geographical resolution, most of the spatial variation in CTR can be accounted for by behavioral histories, leaving little independent role for geographical data. The contrast between county and city levels thus highlights how geographical data's incremental contribution diminishes as spatial units become more granular, consistent with the idea that geographical data largely serves as a proxy for spatial confounding that behavioral data can eventually capture.

table[table omitted — 723 chars of source]

Taken together, these results show that geographical data explains substantial spatial variation in ad responses when behavioral information is excluded, but much of this variation disappears once behavioral data are incorporated. Moreover, the remaining spatial structure becomes smaller at finer spatial resolutions. To examine whether this attenuation also depends on the amount of behavioral information available for each user, we next split impressions into sparse and rich behavioral-history subsets and repeat the analysis.

Residual Spatial Dependence under Sparse versus Rich Behavioral Histories

As discussed earlier, the ability of the behavioral model $f(X^B)$ to account for variation in CTR depends on the amount of behavioral history available for each user. To examine this effect, we split the sample into two groups: (i) a sparse-history group, consisting of the first 50% of impressions observed for each user, and (ii) a rich-history group, consisting of an equal number of later impressions for the same users. This design allows us to compare scenarios where targeting models rely on limited versus extensive behavioral histories.

figure[figure omitted — 322 chars of source]

Figure (ref) illustrates the spatial distribution of behavior-adjusted residuals at the county level for the two groups. In the sparse-history case (left panel), several large contiguous clusters of high residuals remain even after controlling for $f(X^B)$, indicating that location continues to explain a meaningful share of CTR variation. By contrast, in the rich-history case (right panel), residual clustering is visibly weaker and more fragmented, suggesting that much of the spatial dependence disappears once users accumulate longer behavioral records.

table[table omitted — 625 chars of source]

Table (ref) formalizes these visual patterns. Under sparse histories, Moran’s $I$ is positive and statistically significant at the county level, while Geary’s $C$ also indicates non-random clustering. These results confirm that when behavior is limited, geographical data captures residual variation that remains spatially structured. In the rich-history case, however, both statistics decline substantially in magnitude and lose statistical significance across county and city levels. This indicates that once behavioral data are rich, $f(X^B)$ already accounts for nearly all of the systematic spatial correlation in CTR.

The comparison between sparse and rich behavioral histories also sheds light on the mechanisms through which geographical data matters. When behavioral histories are sparse, residual spatial structure may arise from spatial influence, spatial confounding, or both. Observing how this structure evolves as behavioral data becomes richer is therefore informative about which mechanism dominates. Spatial influence is inherently relational and cannot be recovered from individual behavioral data $X^B$, regardless of how rich those histories become. If spatial influence were the primary driver, residual spatial autocorrelation would persist even after conditioning on $X^B$. Instead, we find that spatial autocorrelation largely disappears once behavioral histories are rich. This pattern rules out spatial influence and suggests that spatial confounding is the dominant channel: geographical data initially captures similarities in preferences across nearby users, but these similarities are eventually learned directly from individual behavioral histories, leaving little residual role for $X^G$.

Managerial and Policy Implications

From the perspective of advertising platforms and advertisers, our results highlight a clear trade-off in the use of geographical information for engagement modeling. Geographical data provides incremental value primarily during early cold-start phases, when behavioral histories are sparse and cannot yet support effective personalization. As behavioral data accumulates, it increasingly substitutes for the value previously provided by geographical information, making location-based targeting largely redundant. This pattern suggests that firms can restrict the use of geographical data to early cold-start stages and gradually transition to behavioral targeting as behavioral histories mature. Such an approach allows firms to reduce unnecessary collection of location information while maintaining targeting performance, and may also strengthen consumer trust through transparent, user-controlled opt-out mechanisms for geographical tracking martin2017data.

From a consumer perspective, the substitutability of geographical information by behavioral information implies that consumers can safely opt out of location tracking once sufficient behavioral history has accumulated. In this regime, relevant targeting can be maintained without continued use of geographical information, allowing consumers to limit exposure of sensitive location attributes without materially affecting engagement outcomes.

From a regulatory perspective, the limited incremental value of geographical information in the presence of rich behavioral histories raises questions about the justification for continued large-scale location tracking. While geographical information may serve a temporary role in early cold-start settings, its long-run substitutability suggests that stricter oversight of persistent geographical tracking is warranted, particularly in digital advertising environments where comparable targeting performance can be achieved with behavioral information alone.

Conclusion

In this paper, we develop a unified framework to evaluate the value of information in digital advertising, integrating economic theory, machine learning, and causal inference. Conceptually, we extend decision-based definitions of information value to test whether behavioral and geographical data act as complements or substitutes. Methodologically, we combine LSTM architectures with attention mechanisms to capture temporal and spatial dynamics and use inverse propensity scoring to recover counterfactual performance from observational data. Together, these elements provide a generalizable approach for comparing multiple information sets in high-dimensional environments, with applications extending beyond advertising.

Our results show that both geographical and behavioral data enhance targeting, but their roles change as behavioral histories become richer. At the aggregate level, complementarities and substitutions offset one another, yielding no clear overall interaction. Examining heterogeneity as users accumulate behavioral histories allows us to uncover a clear pattern. Geographical data explains most variation when behavioral information is minimal, complements behavioral data when histories are sparse, and becomes a substitute once histories are rich. This progression indicates that the value of geographical data is temporary and concentrated in the early stages before behavioral information alone suffices for accurate targeting.

Future research could extend our analysis in several ways. First, beyond in-app advertising, the value of geographical data may differ in sectors where location remains central, such as transportation, local retail, or urban planning. Second, future work should trace the privacy–utility frontier by evaluating policy value under explicit data-use constraints, including coarse geolocation, $k$-anonymity, differential privacy, or federated learning. Finally, our analysis focuses on online in-app advertising, where engagement is primarily shaped by behavioral information rather than physical location. We do not generalize these conclusions to industries in which location carries intrinsic value, such as ride-hailing, food delivery, local event promotions, or brick-and-mortar retail discounts. In such contexts, geographical information may remain a critical input, and further research is needed to assess whether behavioral information can fully substitute for location-based targeting.

Competing Interests Declaration

Author(s) have no competing interests to declare.

{\singlespacing

thebibliography{49} \expandafter\ifx\csname urlstyle\endcsname\relax \else \fi \bibitem[Acquisti et al.(2016)Acquisti, Taylor, and Wagman]{acquisti2016economics} Alessandro Acquisti, Curtis Taylor, and Liad Wagman. \newblock The economics of privacy. \newblock Journal of Economic Literature, 54\penalty0 (2):\penalty0 442--492, 2016. \bibitem[Anagnostopoulos et al.(2008)Anagnostopoulos, Kumar, and Mahdian]{anagnostopoulos2008influence} Aris Anagnostopoulos, Ravi Kumar, and Mohammad Mahdian. \newblock Influence and correlation in social networks. \newblock In Proceedings of the 14th ACM SIGKDD international conference on Knowledge discovery and data mining, pages 7--15, 2008. \bibitem[Aridor et al.(2024)Aridor, Che, Hollenbeck, McCarthy, and Kaiser]{aridor2024evaluating} Guy Aridor, Yeon-Koo Che, Brett Hollenbeck, Daniel McCarthy, and Maximilian Kaiser. \newblock Evaluating the impact of privacy regulation on e-commerce firms: Evidence from apple’s app tracking transparency. \newblock 2024. \bibitem[Ascarza(2018)]{ascarza2018retention} Eva Ascarza. \newblock Retention futility: Targeting high-risk customers might be ineffective. \newblock Journal of marketing Research, 55\penalty0 (1):\penalty0 80--98, 2018. \bibitem[Blackwell(1953)]{blackwell1953comparison} David Blackwell. \newblock Equivalent comparisons of experiments. \newblock Annals of Mathematical Statistics, 24\penalty0 (2):\penalty0 265--272, 1953. \bibitem[Bollinger and Gillingham(2012)]{bollinger2012peer} Bryan Bollinger and Kenneth Gillingham. \newblock Peer effects in the diffusion of solar photovoltaic panels. \newblock Marketing Science, 31\penalty0 (6):\penalty0 900--912, 2012. \bibitem[B{\"o}rgers et al.(2013)B{\"o}rgers, Hernando-Veciana, and Kr{\"a}hmer]{borgers2013signals} Tilman B{\"o}rgers, Angel Hernando-Veciana, and Daniel Kr{\"a}hmer. \newblock When are signals complements or substitutes? \newblock Journal of Economic Theory, 148\penalty0 (1):\penalty0 165--195, 2013. \bibitem[Bradlow et al.(2005)Bradlow, Bronnenberg, Russell, Arora, Bell, Duvvuri, Hofstede, Sismeiro, Thomadsen, and Yang]{bradlow2005spatial} Eric T Bradlow, Bart Bronnenberg, Gary J Russell, Neeraj Arora, David R Bell, Sri Devi Duvvuri, Frankel Ter Hofstede, Catarina Sismeiro, Raphael Thomadsen, and Sha Yang. \newblock Spatial models in marketing. \newblock \emph{Marketing letters}, 16\penalty0 (3):\penalty0 267--278, 2005. \bibitem[Bronnenberg and Mahajan(2001)]{bronnenberg2001unobserved} Bart J Bronnenberg and Vijay Mahajan. \newblock Unobserved retailer behavior in multimarket data: Joint spatial dependence in market shares and promotion variables. \newblock \emph{Marketing Science}, 20\penalty0 (3):\penalty0 284--299, 2001. \bibitem[Bronnenberg et al.(2009)Bronnenberg, Dhar, and Dub{\'e}]{bronnenberg2009brand} Bart J Bronnenberg, Sanjay K Dhar, and Jean-Pierre H Dub{\'e}. \newblock Brand history, geography, and the persistence of brand shares. \newblock \emph{Journal of political Economy}, 117\penalty0 (1):\penalty0 87--115, 2009. \bibitem[Cui et al.(2019)Cui, Zhang, and Bassamboo]{cui2019learning} Ruomeng Cui, Dennis J Zhang, and Achal Bassamboo. \newblock Learning from inventory availability information: Evidence from field experiments on amazon. \newblock \emph{Management Science}, 65\penalty0 (3):\penalty0 1216--1235, 2019. \bibitem[De Montjoye et al.(2013)De Montjoye, Hidalgo, Verleysen, and Blondel]{de2013unique} Yves-Alexandre De Montjoye, C{\'e}sar A Hidalgo, Michel Verleysen, and Vincent D Blondel. \newblock Unique in the crowd: The privacy bounds of human mobility. \newblock \emph{Scientific reports}, 3\penalty0 (1):\penalty0 1--5, 2013. \bibitem[De Montjoye et al.(2018)De Montjoye, Gambs, Blondel, Canright, De Cordes, Deletaille, Eng{\o}-Monsen, Garcia-Herranz, Kendall, Kerry, et al.]{de2018privacy} Yves-Alexandre De Montjoye, S{\'e}bastien Gambs, Vincent Blondel, Geoffrey Canright, Nicolas De Cordes, S{\'e}bastien Deletaille, Kenth Eng{\o}-Monsen, Manuel Garcia-Herranz, Jake Kendall, Cameron Kerry, et al. \newblock On the privacy-conscientious use of mobile phone data. \newblock \emph{Scientific data}, 5\penalty0 (1):\penalty0 1--6, 2018. \bibitem[eMarketer(2024)]{emarketer2024mobile} eMarketer. \newblock Mobile advertising 2024: In-app ads drive growth past \$200 billion. \newblock \texttt{https://www.emarketer.com/content/mobile-advertising-2024}, 2024. \newblock Accessed August 2025. \bibitem[Geary(1954)]{geary1954contiguity} Robert C Geary. \newblock The contiguity ratio and statistical mapping. \newblock \emph{The incorporated statistician}, 5\penalty0 (3):\penalty0 115--146, 1954. \bibitem[Ghose et al.(2019)Ghose, Li, and Liu]{ghose2019mobile} Anindya Ghose, Beibei Li, and Siyuan Liu. \newblock Mobile targeting using customer trajectory patterns. \newblock \emph{Management Science}, 65\penalty0 (11):\penalty0 5027--5049, 2019. \bibitem[Goldfarb and Tucker(2011)]{goldfarb2011online} Avi Goldfarb and Catherine Tucker. \newblock Online display advertising: Targeting and obtrusiveness. \newblock \emph{Marketing Science}, 30\penalty0 (3):\penalty0 389--404, 2011. \bibitem[Grbovic and Cheng(2018)]{grbovic2018real} Mihajlo Grbovic and Haibin Cheng. \newblock Real-time personalization using embeddings for search ranking at airbnb. \newblock In \emph{Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining}, pages 311--320, 2018. \bibitem[Guo and Berkhahn(2016)]{guo2016entity} Cheng Guo and Felix Berkhahn. \newblock Entity embeddings of categorical variables. \newblock \emph{arXiv preprint arXiv:1604.06737}, 2016. \bibitem[Hirano et al.(2003)Hirano, Imbens, and Ridder]{hirano2003efficient} Keisuke Hirano, Guido W Imbens, and Geert Ridder. \newblock Efficient estimation of average treatment effects using the estimated propensity score. \newblock \emph{Econometrica}, 71\penalty0 (4):\penalty0 1161--1189, 2003. \bibitem[Hochreiter(1997)]{hochreiter1997long} S Hochreiter. \newblock Long short-term memory. \newblock \emph{Neural Computation MIT-Press}, 1997. \bibitem[Horvitz and Thompson(1952)]{horvitz1952generalization} Daniel G Horvitz and Donovan J Thompson. \newblock A generalization of sampling without replacement from a finite universe. \newblock \emph{Journal of the American statistical Association}, 47\penalty0 (260):\penalty0 663--685, 1952. \bibitem[Iyengar et al.(2011)Iyengar, Van den Bulte, and Valente]{iyengar2011opinion} Raghuram Iyengar, Christophe Van den Bulte, and Thomas W Valente. \newblock Opinion leadership and social contagion in new product diffusion. \newblock \emph{Marketing science}, 30\penalty0 (2):\penalty0 195--212, 2011. \bibitem[Jerath and Miller(2024)]{jerath2024consumers} Kinshuk Jerath and Klaus M Miller. \newblock Consumers' perceived privacy violations in online advertising. \newblock \emph{arXiv preprint arXiv:2403.03612}, 2024. \bibitem[Johnson et al.(2020)Johnson, Shriver, and Du]{johnson2020consumer} Garrett A Johnson, Scott K Shriver, and Shaoyin Du. \newblock Consumer privacy choice in online advertising: Who opts out and at what cost to industry? \newblock \emph{Marketing Science}, 39\penalty0 (1):\penalty0 33--51, 2020. \bibitem[Kallus(2019)]{kallus2019stable} Nathan Kallus. \newblock Intrinsically efficient, stable, and bounded off-policy evaluation for reinforcement learning. \newblock In \emph{Advances in Neural Information Processing Systems}, 2019. \bibitem[Kamenica(2019)]{kamenica2019bayesian} Emir Kamenica. \newblock Bayesian persuasion and information design. \newblock \emph{Annual Review of Economics}, 11\penalty0 (1):\penalty0 249--272, 2019. \bibitem[Kim et al.(2022)Kim, Bradlow, and Iyengar]{kim2022selecting} Mingyung Kim, Eric T Bradlow, and Raghuram Iyengar. \newblock Selecting data granularity and model specification using the scaled power likelihood with multiple weights. \newblock \emph{Marketing Science}, 41\penalty0 (4):\penalty0 848--866, 2022. \bibitem[Larson et al.(2005)Larson, Bradlow, and Fader]{larson2005exploratory} Jeffrey S Larson, Eric T Bradlow, and Peter S Fader. \newblock An exploratory look at supermarket shopping paths. \newblock \emph{International Journal of research in Marketing}, 22\penalty0 (4):\penalty0 395--414, 2005. \bibitem[Loshchilov and Hutter(2019)]{Loshchilov2019Decoupled} Ilya Loshchilov and Frank Hutter. \newblock Decoupled weight decay regularization. \newblock \emph{International Conference on Learning Representations (ICLR)}, 2019. \bibitem[Luo and Ranjan(2025)]{luo2025mapping} Bowen Luo and Bhoomija Ranjan. \newblock Mapping spatial heterogeneity in retail advertising effectiveness. \newblock \emph{Journal of Marketing Research}, 62\penalty0 (6):\penalty0 1063--1080, 2025. \bibitem[Ma et al.(2015)Ma, Krishnan, and Montgomery]{ma2015latent} Liye Ma, Ramayya Krishnan, and Alan L Montgomery. \newblock Latent homophily or social influence? an empirical analysis of purchase within a social network. \newblock \emph{Management Science}, 61\penalty0 (2):\penalty0 454--473, 2015. \bibitem[Manchanda et al.(2008)Manchanda, Xie, and Youn]{manchanda2008role} Puneet Manchanda, Ying Xie, and Nara Youn. \newblock The role of targeted communication and contagion in product adoption. \newblock \emph{Marketing Science}, 27\penalty0 (6):\penalty0 961--976, 2008. \bibitem[Martin et al.(2017)Martin, Borah, and Palmatier]{martin2017data} Kelly D Martin, Abhishek Borah, and Robert W Palmatier. \newblock Data privacy: Effects on customer and firm performance. \newblock \emph{Journal of marketing}, 81\penalty0 (1):\penalty0 36--58, 2017. \bibitem[McCaffrey et al.(2013)McCaffrey, Griffin, Almirall, Slaughter, Ramchand, and Burgette]{mccaffrey2013tutorial} Daniel F McCaffrey, Beth Ann Griffin, Daniel Almirall, Mary Ellen Slaughter, Rajeev Ramchand, and Lane F Burgette. \newblock A tutorial on propensity score estimation for multiple treatments using generalized boosted models. \newblock \emph{Statistics in medicine}, 32\penalty0 (19):\penalty0 3388--3414, 2013. \bibitem[Mirrokni et al.(2010)Mirrokni, Muthukrishnan, and Nadav]{mirrokni2010quasi} Vahab Mirrokni, S Muthukrishnan, and Uri Nadav. \newblock Quasi-proportional mechanisms: Prior-free revenue maximization. \newblock In \emph{Latin American Symposium on Theoretical Informatics}, pages 565--576. Springer, 2010. \bibitem[Moran(1950)]{moran1950notes} Patrick AP Moran. \newblock Notes on continuous stochastic phenomena. \newblock \emph{Biometrika}, 37\penalty0 (1/2):\penalty0 17--23, 1950. \bibitem[Narang and Luco(2025)]{narang2025privacy} Unnati Narang and Fernando Luco. \newblock Privacy and prediction: how useful are geo-tracking data for predicting consumer visits? \newblock \emph{Quantitative Marketing and Economics}, 23\penalty0 (4):\penalty0 523--544, 2025. \bibitem[Pew-Center(2024)]{pew2024smartphone} Pew-Center. \newblock Mobile fact sheet. \newblock \texttt{https://www.pewresearch.org/internet/fact-sheet/mobile/}, 2024. \newblock Accessed August 2025. \bibitem[Quadrana et al.(2018)Quadrana, Cremonesi, and Jannach]{quadrana2018sequence} Massimo Quadrana, Paolo Cremonesi, and Dietmar Jannach. \newblock Sequence-aware recommender systems. \newblock \emph{ACM computing surveys (CSUR)}, 51\penalty0 (4):\penalty0 1--36, 2018. \bibitem[Rafieian(2023)]{rafieian2023optimizing} Omid Rafieian. \newblock Optimizing user engagement through adaptive ad sequencing. \newblock \emph{Marketing Science}, 42\penalty0 (5):\penalty0 910--933, 2023. \bibitem[Rafieian and Yoganarasimhan(2021)]{Rafieian2021} Omid Rafieian and Hema Yoganarasimhan. \newblock Targeting and privacy in mobile advertising. \newblock \emph{Marketing Science}, 40\penalty0 (2):\penalty0 193--218, 2021. \bibitem[Rafieian and Yoganarasimhan(2023)]{rafieian2023ai} Omid Rafieian and Hema Yoganarasimhan. \newblock Ai and personalization. \newblock \emph{Artificial Intelligence in Marketing}, pages 77--102, 2023. \bibitem[Rossi et al.(1996)Rossi, McCulloch, and Allenby]{rossi1996value} Peter E Rossi, Robert E McCulloch, and Greg M Allenby. \newblock The value of purchase history data in target marketing. \newblock \emph{Marketing Science}, 15\penalty0 (4):\penalty0 321--340, 1996. \bibitem[Rumelhart et al.(1986)Rumelhart, Hinton, and Williams]{rumelhart1986learning} David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams. \newblock Learning representations by back-propagating errors. \newblock \emph{nature}, 323\penalty0 (6088):\penalty0 533--536, 1986. \bibitem[Smith et al.(2023)Smith, Seiler, and Aggarwal]{smith2023optimal} Adam N Smith, Stephan Seiler, and Ishant Aggarwal. \newblock Optimal price targeting. \newblock \emph{Marketing Science}, 42\penalty0 (3):\penalty0 476--499, 2023. \bibitem[Tucker(2014)]{tucker2014social} Catherine E Tucker. \newblock Social networks, personalized advertising, and privacy controls. \newblock \emph{Journal of marketing research}, 51\penalty0 (5):\penalty0 546--562, 2014. \bibitem[Vaswani et al.(2017)Vaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser, and Polosukhin]{vaswani2017attention} Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, {\L}ukasz Kaiser, and Illia Polosukhin. \newblock Attention is all you need. \newblock \emph{Advances in neural information processing systems}, 30, 2017. \bibitem[Wernerfelt et al.(2025)Wernerfelt, Tuchman, Shapiro, and Moakler]{wernerfelt2025estimating} Nils Wernerfelt, Anna Tuchman, Bradley T Shapiro, and Robert Moakler. \newblock Estimating the value of offsite tracking data to advertisers: Evidence from meta. \newblock \emph{Marketing Science}, 44\penalty0 (2):\penalty0 268--286, 2025.

}

\setcounter{table}{0} \setcounter{page}{0} \setcounter{figure}{0}