EconBase
← All papers

Estimating Causal Effects from Data Generated by Stochastic Algorithms

Susan Athey, Guido Imbens, Zoe Ji

arXiv 7 Jul 2026 · Statistics — Methodology

arXiv:2607.05792 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Recommendation systems and chatbots present content to users, typically using stochastic algorithms that select the content based on user characteristics or context. Examples of content include chat responses, videos, or items available for purchase. Scientists and application developers are often interested in whether characteristics of content increase outcomes such as user engagement. Estimates of such causal effects may guide content providers to generate content that emphasize desirable features. However, in settings with a large content library or where content is generated uniquely for a given user, it can be difficult to use observational data to learn the causal effect of content features, because the content a user sees is tailored to that user, and because content varies in many dimensions. This paper proposes a new method for estimating the impact of content features using observational data, when the algorithm that determines user exposure incorporates some randomization, and when two additional data elements are logged for each user: $(i)$ the identity of at least one item that could have been exposed to the user, but was not (the unexposed item); $(ii)$ an estimate of the ratio of the probability that the unexposed item would have been shown to the probability that the exposed item was shown. We show that causal effects of features are identified in this setting, even in the presence of unobserved confounders that affect both user preferences and the identity of the considered pair of items (exposed and unexposed). Our estimator differs from prior approaches in terms of what data is used and how the estimator is constructed.

Citation extraction

56
references
67
in-text mentions
56
distinct cited
13
self-citations
19,273
main-text words

appendix boundary found by appendix_command · 95% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Jason Hartford, Greg Lewis, Kevin Leyton-Brown, and Matt Taddy (2017) Deep iv: A flexible approach for counterfactual prediction0.84333100%
2Ruining He and Julian McAuley (2016) Vbpr: visual bayesian personalized ranking from implicit feedback0.84333100%
3Léon Bottou, Jonas Peters, Joaquin Quiñonero-Candela, Denis X. Charl… (2013) Counterfactual reasoning and learning systems: The example of computational advertising0.64422100%
4Thorsten Joachims, Adith Swaminathan, and Tobias Schnabel (2017) Unbiased learning-to-rank with biased feedback0.64422100%
5Lihong Li, Wei Chu, John Langford, and Robert E. Schapire (2010) A contextual-bandit approach to personalized news article recommendation0.64422100%
6Tobias Schnabel, Adith Swaminathan, Peter I Frazier, and Thorsten Jo… (2016) Unbiased comparative evaluation of ranking functions0.64422100%
7Masatoshi Uehara, Chengchun Shi, and Nathan Kallus (2022) A review of off-policy evaluation in reinforcement learning0.64422100%
8Tyler J VanderWeele and Miguel A Hernan (2013) Causal inference under multiple versions of treatment0.64422100%
9Guido W Imbens and Donald B Rubin (2015) Causal Inference in Statistics, Social, and Biomedical Sciences self0.51121100%
10Keshav Agrawal, Susan Athey, Ayush Kanodia, Shanjukta Nath, and Emil… (2026) The economics of algorithmic personalization: Evidence from an educational technology platform self0.40511100%

Showing the top 10 of 56 scored citations.