EconBase
← All papers

The Challenge of Using LLMs to Simulate Human Behavior: A Causal Inference Perspective

George Gui, Olivier Toubia

arXiv 24 Dec 2023 · Artificial Intelligence · 37 citations (OpenAlex)

arXiv:2312.15524 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Large Language Models (LLMs) have shown impressive potential to simulate human behavior. We identify a fundamental challenge in using them to simulate experiments: when LLM-simulated subjects are blind to the experimental design (as is standard practice with human subjects), variations in treatment systematically affect unspecified variables that should remain constant, violating the unconfoundedness assumption. Using demand estimation as a context and an actual experiment as a benchmark, we show this can lead to implausible results. While confounding may in principle be addressed by controlling for covariates, this can compromise ecological validity in the context of LLM simulations: controlled covariates become artificially salient in the simulated decision process, which introduces focalism. This trade-off between unconfoundedness and ecological validity is usually absent in traditional experimental design and represents a unique challenge in LLM simulations. We formalize this challenge theoretically, showing it stems from ambiguous prompting strategies, and hence cannot be fully addressed by improving training data or by fine-tuning. Alternative approaches that unblind the experimental design to the LLM show promise. Our findings suggest that effectively leveraging LLMs for experimental simulations requires fundamentally rethinking established experimental design practices rather than simply adapting protocols developed for human subjects.

Citation extraction

35
references
53
in-text mentions
35
distinct cited
1
self-citations
7,436
main-text words

appendix boundary found by appendix_command · 64% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Toubia, Olivier and Gui, George Z and Peng, Tianyi and Merlau, Danie… (2025) Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions self1.00073100%
2Argyle, Lisa P. and Busby, Ethan C. and Fulda, Nancy and Gubler, Jos… (2023) Out of One, Many: Using Language Models to Simulate Human Samples0.92843100%
3Hewitt, Luke and Ashokkumar, Ashwini and Ghezae, Isaias and Willer,… (2024) Predicting results of social science experiments using large language models0.81142100%
4Horton, John J (2023) Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus?0.73732100%
5Brand, James and Israeli, Ayelet and Ngwe, Donald (2023) Using GPT for Market Research0.64422100%
6Berke, Alex and Calacci, Dan and Mahari, Robert and Yabe, Takahiro a… (2024) Open e-commerce 1.0, five years of crowdsourced US Amazon purchase histories with user demographics0.5112250%
7DellaVigna, Stefano and Gentzkow, Matthew (2019) Uniform Pricing in U.S. Retail Chains*0.5112250%
8Ludwig, Jens and Mullainathan, Sendhil and Rambachan, Ashesh (2024) Large language models: An applied econometric framework0.51121100%
9Angrist, Joshua D. and Pischke, Jörn-Steffen (2009) Mostly Harmless Econometrics: An Empiricist's Companion0.40511100%
10Imbens, Guido W. and Rubin, Donald B (2015) Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction0.40511100%

Showing the top 10 of 35 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions0.85584
2Statistical Foundations of LLM-based A/B Testing: A Surrogacy Framework for Human Causal Inference0.51121
3Leveraging LLMs to Improve Experimental Design: A Generative Stratification Approach0.40511
4LLM Personas as a Substitute for Field Experiments in Method Benchmarking0.40511
5Tabular Foundation Models for Discrete Choice Estimation0.40511