EconBase
← All papers

Causal Inference on Outcomes Learned from Text

Iman Modarressi, Jann Spiess, Amar Venugopal

arXiv 2 Mar 2025 · Econometrics

arXiv:2503.00725 · PDF · Extracted main text

Abstract

We propose a machine-learning tool that yields causal inference on text in randomized trials. Based on a simple econometric framework in which text may capture outcomes of interest, our procedure addresses three questions: First, is the text affected by the treatment? Second, which outcomes is the effect on? And third, how complete is our description of causal effects? To answer all three questions, our approach uses large language models (LLMs) that suggest systematic differences across two groups of text documents and then provides valid inference based on costly validation. Specifically, we highlight the need for sample splitting to allow for statistical validation of LLM outputs, as well as the need for human labeling to validate substantive claims about how documents differ across groups. We illustrate the tool in a proof-of-concept application using abstracts of academic manuscripts.

Citation extraction

46
references
75
in-text mentions
46
distinct cited
0
self-citations
13,596
main-text words

appendix boundary found by appendix_command · 80% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Zhong, Ruiqi, Peter Zhang, Steve Li, Jinwoo Ahn, Dan Klein, and Jaco… (2023) Goal driven discovery of distributional differences via language descriptions1.00054100%
2Ludwig, Jens, Sendhil Mullainathan, and Jann Spiess (2017) Machine-Learning Tests for Effects on Multiple Outcomes0.92844100%
3Egami, Naoki, Christian J Fong, Justin Grimmer, Margaret E Roberts,… (2022) How to make causal inferences using texts0.92844100%
4Angelopoulos, Anastasios N, Stephen Bates, Clara Fannjiang, Michael… (2023) Prediction-powered inference0.92843100%
5Ludwig, Jens, Sendhil Mullainathan, and Ashesh Rambachan (2024) Large language models: An applied econometric framework0.92843100%
6Zhong, Ruiqi, Charlie Snell, Dan Klein, and Jacob Steinhardt (2022) Describing differences between text distributions with natural language0.92843100%
7Fudenberg, Drew, Jon Kleinberg, Annie Liang, and Sendhil Mullainathan (2022) Measuring the completeness of economic models0.84333100%
8Ludwig, Jens and Sendhil Mullainathan (2024) Machine learning as a tool for hypothesis generation0.84333100%
9Egami, Naoki, Musashi Hinck, Brandon M Stewart, and Hanying Wei (2024) Using large language model annotations for the social sciences: A general framework of using predicted variables in downstream a…0.73732100%
10Robins, James M, Andrea Rotnitzky, and Lue Ping Zhao (1994) Estimation of regression coefficients when some regressors are not always observed0.73732100%

Showing the top 10 of 46 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach0.95073
2Participation and Representation in Local Government Speech0.64422
3Inference for Regression with Variables Generated by AI or Machine Learning0.40511
4Large Language Models: An Applied Econometric Framework0.40511
5Causal Effect Estimation with Latent Textual Treatments0.40511
6Econometrics with Pre-Trained Embeddings for Unstructured Data0.40511