arXiv 7 Jul 2025 · Statistics — Methodology
arXiv:2507.05175 · PDF · DOI · OpenAlex · Extracted main text
Major advertising platforms recently increased privacy protections by limiting advertisers' access to individual-level data. Instead of providing access to granular raw data, the platforms only allow a limited number of aggregate queries to a dataset, which is further protected by adding differentially private noise. This paper studies whether and how advertisers can design effective targeting policies within these restrictive privacy preserving data environments. To achieve this, I develop a probabilistic machine learning method based on Bayesian optimization, which facilitates dynamic data exploration. Since Bayesian optimization was designed to sample points from a function to find its maximum, it is not applicable to aggregate queries and to targeting. Therefore, I introduce two innovations: (i) integral updating of posteriors which allows to select the best regions of the data to query rather than individual points and (ii) a targeting-aware acquisition function that dynamically selects the most informative regions for the targeting task. I identify the conditions of the dataset and privacy environment that necessitate the use of such a "smart" querying strategy. I apply the strategic querying method to the Criteo AI Labs dataset for uplift modeling (Diemert et al., 2018) that contains visit and conversion data from 14M users. I show that an intuitive benchmark strategy only achieves 33% of the non-privacy-preserving targeting potential in some cases, while my strategic querying method achieves 97-101% of that potential, and is statistically indistinguishable from Causal Forest (Athey et al., 2019): a state-of-the-art non-privacy-preserving machine learning targeting method.
appendix boundary found by appendix_command · 98% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Athey, S., Tibshirani, J., and Wager, S (2019) Generalized random forests | 1.000 | 5 | 4 | 100% |
| 2 | Smith, M. T., Álvarez, M. A., and Lawrence, N. D (2018) Gaussian process regression for binned data | 0.928 | 4 | 3 | 100% |
| 3 | Diemert, E., Betlei, A., Renaudin, C., and Amini, M.-R (2018) A large scale benchmark for uplift modeling | 0.843 | 3 | 3 | 100% |
| 4 | Dew, R (2024) Adaptive preference measurement with unstructured data | 0.737 | 3 | 3 | 67% |
| 5 | Athey, S., Catalini, C., and Tucker, C (2017) The digital privacy paradox: Small money, small costs, small talk | 0.644 | 2 | 2 | 100% |
| 6 | Chen, Z., Karthik, P., Chee, Y. M., and Tan, V. Y (2024) Fixed-budget differentially private best arm identification | 0.644 | 2 | 2 | 100% |
| 7 | Kushner, H. J (1964) A new method of locating the maximum point of an arbitrary multipeak curve in the presence of noise | 0.644 | 2 | 2 | 100% |
| 8 | Shchetkina, A. and Berman, R (2024) When is heterogeneity actionable for personalization? self | 0.644 | 2 | 2 | 100% |
| 9 | Snoek, J., Larochelle, H., and Adams, R. P (2012) Practical Bayesian optimization of machine learning algorithms | 0.644 | 2 | 2 | 100% |
| 10 | Wernerfelt, N., Tuchman, A., Shapiro, B. T., and Moakler, R (2025) Estimating the value of offsite tracking data to advertisers: Evidence from Meta | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 67 scored citations.