Olivier Toubia, George Z. Gui, Tianyi Peng, Daniel J. Merlau, Ang Li, Haozhe Chen
arXiv 23 May 2025 · Computers and Society · 1 citations (OpenAlex)
arXiv:2505.17479 · PDF · DOI · OpenAlex · Extracted main text
LLM-based digital twin simulation, where large language models are used to emulate individual human behavior, holds great promise for research in AI, social science, and digital experimentation. However, progress in this area has been hindered by the scarcity of real, individual-level datasets that are both large and publicly available. This lack of high-quality ground truth limits both the development and validation of digital twin methodologies. To address this gap, we introduce a large-scale, public dataset designed to capture a rich and holistic view of individual human behavior. We survey a representative sample of $N = 2,058$ participants (average 2.42 hours per person) in the US across four waves with 500 questions in total, covering a comprehensive battery of demographic, psychological, economic, personality, and cognitive measures, as well as replications of behavioral economics experiments and a pricing survey. The final wave repeats tasks from earlier waves to establish a test-retest accuracy baseline. Initial analyses suggest the data are of high quality and show promise for constructing digital twins that predict human behavior well at the individual and aggregate levels. By making the full dataset publicly available, we aim to establish a valuable testbed for the development and benchmarking of LLM-based persona simulations. Beyond LLM applications, due to its unique breadth and scale the dataset also enables broad social science research, including studies of cross-construct correlations and heterogeneous treatment effects.
appendix boundary found by appendix_command · 24% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Gui, G. and Toubia, O (2023) The challenge of using llms to simulate human behavior: A causal inference perspective self | 0.855 | 8 | 4 | 62% |
| 2 | Kahneman, D. and Tversky, A (1973) On the psychology of prediction | 0.843 | 4 | 4 | 75% |
| 3 | Dean, M. and Ortoleva, P (2019) The empirical relationship between nonstandard economic behaviors | 0.825 | 16 | 3 | 56% |
| 4 | Stanovich, K. E. and West, R. F (2008) On the relative independence of thinking biases and cognitive ability | 0.805 | 46 | 5 | 52% |
| 5 | Baron, J. and Hershey, J. C (1988) Outcome bias in decision evaluation | 0.794 | 6 | 3 | 50% |
| 6 | Furnas, A. C. and LaPira, T. M (2024) The people think what i think: False consensus and unelected elite misperception of public opinion | 0.737 | 4 | 3 | 50% |
| 7 | Tversky, A. and Kahneman, D (1981) The framing of decisions and the psychology of choice | 0.737 | 4 | 3 | 50% |
| 8 | Tversky, A. and Kahneman, D (1983) Extensional versus intuitive reasoning: The conjunction fallacy in probability judgment | 0.737 | 4 | 3 | 50% |
| 9 | Epley, N., Keysar, B., Van Boven, L., and Gilovich, T (2004) Perspective taking as egocentric anchoring and adjustment | 0.737 | 3 | 3 | 67% |
| 10 | Park, J. S., Zou, C. Q., Shaw, A., Hill, B. M., Cai, C., Morris, M.… (2024) Generative agent simulations of 1,000 people | 0.737 | 3 | 2 | 100% |
Showing the top 10 of 51 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | The Challenge of Using LLMs to Simulate Human Behavior: A Causal Inference Perspective | 1.000 | 7 | 3 |