EconBase
← Back to paper

Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

39,782 characters · 8 sections · 91 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions

abstractLLM-based digital twin simulation, where large language models are used to emulate individual human behavior, holds great promise for research in AI, social science, and digital experimentation. However, progress in this area has been hindered by the scarcity of real, individual-level datasets that are both large and publicly available. This lack of high-quality ground truth limits both the development and validation of digital twin methodologies. To address this gap, we introduce a large-scale, public dataset designed to capture a rich and holistic view of individual human behavior. We survey a representative sample of $N = 2,058$ participants (average 2.42 hours per person) in the US across four waves with 500 questions in total, covering a comprehensive battery of demographic, psychological, economic, personality, and cognitive measures, as well as replications of behavioral economics experiments and a pricing survey. The final wave repeats tasks from earlier waves to establish a test-retest accuracy baseline. Initial analyses suggest the data are of high quality and show promise for constructing digital twins that predict human behavior well at the individual and aggregate levels. By making the full dataset publicly available, we aim to establish a valuable testbed for the development and benchmarking of LLM-based persona simulations. Beyond LLM applications, due to its unique breadth and scale the dataset also enables broad social science research, including studies of cross-construct correlations and heterogeneous treatment effects. \footnote{The dataset is publicly available at \url{https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500/}, and the LLM simulation code can be found at \url{https://github.com/tianyipeng-lab/Digital-Twin-Simulation}. }

Introduction

The rise of large language models (LLMs) like GPT has sparked interest across disciplines (including computer science, marketing, economics, psychology and political science) in leveraging these tools to create “silicon samples” which may replicate how these humans would behave in response to any stimuli aher_using_2023,argyle_out_2023, brand_using_2023,dillion_can_2023,hewitt2024predicting,horton_large_2023,li2025llm,park_generative_2023,qin2024aiturk. If these LLM-simulations can be a faithful substitute for eliciting responses from their human counterparts, the implications for both academics and practitioners are substantial. Academics could use silicon samples for pilot experiments to pinpoint stimuli with significant impact, thus improving the efficiency of theory development and experimental design. Firms could leverage these realistic simulations to explore different ideas and strategies, thereby improving customer insight and product development. Accordingly, in the recent past we have witnessed a large influx of firms offering services leveraging silicon samples for customer insights (e.g., \href{https://www.syntheticusers.com/}{Synthetic Users}, \href{https://outset.ai/}{Outset AI}, \href{https://www.nexxt.in/}{Nexxt}, \href{https://www.voxpopme.com/}{Voxpopme}, \href{https://www.evidenza.ai/}{Evidenza}, \href{https://www.expectedparrot.com/}{Expected Parrot}, \href{https://meaningful.app/}{Meaningful}, \href{https://xpolls.ai/}{xPolls}, \href{https://www.ipsos.com/en-uk/large-language-models-market-research}{Ipsos}, \href{https://civicsync.com/}{CivicSync}).

While silicon samples may be generated using only demographic information or hypothetical “life stories,” a promising approach consists in creating silicon samples that are “digital twins” of real people. Notably, park2024generative use LLMs to create digital twins of over 1,000 individuals based on transcripts from qualitative interviews, and find that the simulated agents replicated the human participants' responses on the General Social Survey 85% as accurately as participants replicate their own answers two weeks later.

Despite the promise and excitement surrounding digital twins, some uncertainty remains. For example, brucks2023prompt show that the answers provided by LLMs may be overly influenced by the architecture of the prompt, such as the labeling or ordering of options in multiple choice questions. gui2023challenge show that leveraging LLMs to simulate experiments may introduce unwanted confounding, due to the difficulty of clearly instructing the LLM how to draw variables not specified in the prompt. Other research santurkar_whose_2023,motoki2024more,li2025llm suggests LLMs tend to express opinions that are not representative of the (human) population.

Given this background, it is crucial for the academic and practitioner community to validate digital twins in a transparent, reliable, and replicable manner. However, existing datasets present significant limitations that hinder their effectiveness for this purpose. Some datasets are publicly available (e.g., alattar2018introduction,pew_research_2023,santurkar2023whose), but they are not well suited for testing the validity of digital twins because they do not contain behavioral data (e.g., experiments) nor a test-retest accuracy benchmark. Other datasets have this feature, but they are not publicly available (e.g., park2024generative).

In sum, there is no publicly available dataset that combines rich psychological profiles, behavioral data, and demographics from a large, representative sample. As a result, researchers often rely on synthetic or proprietary data, which undermines transparency, reliability, and replicability.

To address this gap, we assemble and publicly share an extensive dataset for a representative sample of $N=2,058$ people who each answered over 500 questions covering a wide range of demographic questions, psychological scales, cognitive performance questions, economic preferences questions, as well as replications of a wide range of within- and between-subject experiments on heuristics and biases taken from the behavioral economics literature. The data was collected across 4 waves of studies lasting on average 2.42 hour per participant in total. Table (ref) gives an overview of the measures collected in each wave and Figure (ref) illustrates our overall approach.

We use the responses to the heuristics and biases questions from waves 1-3 as holdout data, and train the digital twins based on the rest of the data from waves 1-3. Wave 4 repeated the heuristics and biases experiments, providing us with a measure of test-retest accuracy. Future uses of the data may keep the same split, or combine all the data from the 4 waves to create digital twins.

We report encouraging results regarding the quality of the data: correlations between measures have good face validity, we replicate almost all known results from the behavioral economics literature, and the test-retest accuracy is robust.

We also report initial tests of the predictive validity of digital twins constructed using the data. At the individual level, we compute the accuracy of the digital twin predictions on holdout questions, against the test-retest accuracy benchmark as well as a random benchmark. At the aggregate level, we test whether the digital twin simulations replicate the average treatment effects observed in human data. Throughout this process, we explore the type of behavior that can be predicted with higher vs. lower accuracy by the digital twins, to develop insight into the range of potential applications.

To the best of our knowledge, this dataset is unique in its scale and breadth. As such, there is value in the raw results, irrespective of the application to digital twins. We report descriptive statistics and correlations between the dozens of measures we collected. We encourage others to explore heterogeneous treatment effects dean2019empirical,stanovich2008relative.

figure[figure omitted — 175 chars of source]

Methods

We assembled a wide-range of measures proposed in the social science literature over the past several decades. In addition to 14 demographic questions, we included 19 personality tests that measured 26 constructs over 279 questions, 11 cognitive ability tests (85 questions, 11 measures), 10 economic preferences tests (34 questions, 10 measures). We also replicated 11 between-subject experiments (16 questions) and 5 within-subject experiments (32 questions) from the behavioral economics literature. Finally, we administered the pricing study from gui2023challenge, which asks participants to make purchase decisions about 40 different products at randomly selected prices. In total, participants answered 500 questions across the first three waves. Wave 4 repeated the within- and between-subject heuristics and biases experiments from the first three waves as well as the pricing study from wave 3 (88 questions in total). Participants were assigned to the exact same condition in wave 4 as they were in waves 1-3 for each of these experiments, providing us with a clean measure of test-retest accuracy. We programmed the studies on Qualtrics, doing our best to replicate the stimuli and measures from the original papers.\footnote{We made minor adjustments to reflect cultural and societal changes (e.g., in the mental accounting scenarios from thaler1985mental we replaced “Mr. A” and “Mr. B” with “Person A” and “Person B,” and in the sunk cost experiment of stanovich2008relative we replaced video rental stores with coffee shops).} Our study was approved as IRB Protocol (blinded for anonymity).\footnote{Our study was determined as exempt on the basis of “involving benign behavioral interventions in conjunction with the collection of information from an adult subject through verbal or written responses.”}

We launched Wave 1 on Prolific on 01/29/2025, targeting 2,500 representative US respondents (sampled by age, sex, and ethnicity). Participants received \$7 for completing Wave 1.\footnote{We pre-tested each wave to estimate response time and adjusted compensation accordingly.} They were informed that this was the first of four waves and would earn a \$10 bonus for completing all waves (with comprehension checks to ensure understanding). We received 2,509 complete responses. The following week (02/04/2025), we invited these 2,509 participants to Wave 2 (\$7), receiving 2,263 responses. The next week (02/11/2025), we invited them again for Wave 3 (\$7), receiving 2,252 responses. Waves 2 and 3 were closed the next week. On 02/25/2025, we invited the 2,154 participants who had completed Waves 1–3 to Wave 4 (\$6; two-week delay since Wave 4 repeated previous measures), receiving 2,058 responses and closing the wave after one week. Those completing all four waves received an additional \$10 bonus, totaling \$37.

These 2,058 participants who completed all four waves constitute our final sample. Among our final sample, the average response time was 43.88 minutes for wave 1 (std=19.27), 45.31 minutes for wave 2 (std=19.24), 32.66 minutes for wave 3 (std=15.68), and 24.09 minutes for wave 4 (std=12.51). The average total time across all 4 waves was 145.47 minutes (std=56.10).

landscape\tiny \begin{longtable}{p{6.5cm}p{2.5cm}p{8cm}p{1cm}} \caption{Complete list of questions and related measures} \\ \toprule Task (source) & \#Questions (format) & Extracted measure(s) & Wave(s) \\ \midrule \multicolumn{4}{l}{Demographics} \\ \midrule Demographics santurkar_whose_2023& 12 (multiple choice) & region, sex, age, education, race, citizenship, marital status, religion, religious attendance, political party, household income, political ideology (categorical) & 1 \\ \midrule Additional demographics & 2 (multiple choice) & household size, employment status (categorical) & 1 \\ \midrule \multicolumn{3}{l}{Personality Traits} \\ \midrule Big 5 personality test john1999big & 44 (5-point Likert) & extraversion, agreeableness, conscientiousness, neuroticism, openness scores (numerical) & 1 \\ \midrule Need for cognition scale cacioppo1984efficient & 18 (5-point Likert) & need for cognition score (numerical) & 1 \\ \midrule Agentic vs. Communal Values scale trapnell2012agentic & 24 (9-point Likert) & agency score, communion score (numerical) & 1 \\ \midrule Consumer Minimalism scale wilson2022consumer & 12 (5-point Likert) & minimalism score (numerical) & 1 \\ \midrule Empathy scale carre2013basic & 20 (5-point Likert) & basic empathy score (numerical) & 1 \\ \midrule Green values scale haws2014seeing & 6 (5-point Likert) & green score (numerical) & 1 \\ \midrule Social Desirability scale reynolds1982development & 13 (binary choice) & social desirability score (numerical) & 2 \\ \midrule Conscientiousness scale johnson2019s & 8 (9-point Likert) & conscientiousness score (wave 2) (numerical) & 2 \\ \midrule Anxiety scale beck1988inventory & 21 (4-point Likert) & anxiety score (numerical) & 2 \\ \midrule Individualism vs. Collectivism scale triandis1998converging & 16 (5-point Likert) & horizontal/vertical individualism, horizontal/vertical collectivism scores (numerical) & 2 \\ \midrule Selves questionnaire higgins1985self& 3 (open-ended)& n/a & 2 \\ \midrule Regulatory Focus scale fellner2007regulatory& 10 (7-point Likert) &regulatory focus score (numerical)&3 \\ \midrule Tightwads vs. Spendthrift scale rick2008tightwads& 4 (multiple choice) &tightwads vs. spendthrift score (numerical)&3 \\ \midrule Depression scale date1987beck&22 (multiple choice)&depression score (numerical)&3 \\ \midrule Need for uniqueness scale ruvio2008consumers& 12 (5-point Likert) &need for uniqueness score (numerical)&3 \\ \midrule Self-monitoring scale lennox1984revision&13 (6-point Likert)&self-monitoring score (numerical) &3 \\ \midrule Self-concept clarity scale campbell1996self&12 (5-point Likert)& self-concept clarity score (numerical)& 3 \\ \midrule Need for closure scale roets2011item& 15 (5-point Likert) &need for closure score (numerical)&3 \\ \midrule Maximization scale nenkov2008short & 6 (5-point Likert) &maximization score (numerical) &3\\ \midrule \multicolumn{3}{l}{\emph{Cognitive Abilities}} \\ \midrule Cognitive Reflection Test krefeld2024exposing& 4 (open-ended) & CRT score (numerical) & 1 \\ \midrule Fluid intelligence test krefeld2024exposing & 6 (multiple choice) & fluid intelligence score (numerical) & 1 \\ \midrule Crystallized intelligence test krefeld2024exposing & 20 (multiple choice) & crystallized intelligence score (numerical) & 1 \\ \midrule Syllogisms test markovits1989belief & 12 (multiple choice) & syllogism score (numerical) & 1 \\ \midrule Overconfidence dean2019empirical& 1 (numerical) &overconfidence score (own predicted-actual score)& 1\\ \midrule Overplacement dean2019empirical& 1 (numerical) & overplacement score (own predicted score-predicted average)& 1\\ \midrule Financial literacy test johnson2019s& 7 (mult. choice)+1 (num.)&financial literacy score (numerical)&2 \\ \midrule Numeracy test johnson2019s& 8 (numerical)&numeracy score (numerical) &2 \\ \midrule Deductive certainty of Modus Ponens test stanovich2008relative& 4 (binary choice) &deductive certainty score& 2\\ \midrule Forward Flow (free associations) gray2019forward& 20 (open-ended) &forward flow score (average pairwise semantic distance)& 2\\ \midrule Wason Selection Task klauer2007abstract& 1 (multiple choice) &Wason Selection Task score (numerical) &3 \\ \midrule \multicolumn{3}{l}{\emph{Economic Preferences}} \\ \midrule Ultimatum game (sender) guth1982experimental & 1 (multiple choice) & ultimatum-send (percentage sent) & 1 \\ \midrule Ultimatum game (receiver) guth1982experimental & 6 (binary choice) & ultimatum-receive (acceptance probability) & 1 \\ \midrule Mental accounting thaler1985mental& 4 (binary choice) & mental accounting score (% choices consistent with mental account predictions) & 1 \\ \midrule Discount dean2019empirical & 3 (multiple price list) & discount rate (numerical) & 2 \\ \midrule Present bias dean2019empirical & 3 (multiple price list) & present bias (numerical) & 2 \\ \midrule Risk Aversion dean2019empirical & 3 (uncertainty equivalence) & risk aversion coefficient (numerical) & 2 \\ \midrule Loss Aversion dean2019empirical & 4 (uncertainty equivalence) & loss aversion coefficient (numerical) & 2 \\ \midrule Trust game (sender) dean2019empirical & 1 (multiple choice) & trust-send (percentage sent) & 2 \\ \midrule Trust game (receiver) dean2019empirical & 5 (multiple choice) & Trust-return (average percentage returned) & 2 \\ \midrule Trust game (sender) thought listing & 1 (open-ended) &n/a& 2\\ \midrule Trust game (receiver) thought listing & 1 (open-ended) &n/a& 2\\ \midrule Dictator game baron1988outcome& 1 (multiple choice) & dictator-send (percentage sent)& 3\\ \midrule Dictator game thought listing &1 (open-ended) &n/a &3\\ \midrule \multicolumn{3}{l}{\emph{Heuristics and Biases (between subject)}} \\ \midrule Base rate problem kahneman1973psychology& 1 (slider scale) &average prob. assessment (numerical) in each condition (base rate of 30 vs. 70 engineers) &1, 4\\ \midrule Outcome bias baron1988outcome&1 (7-point Likert) &average correctness assessment (numerical) in each condition (success vs. failure)& 1, 4 \\ \midrule Sunk cost fallacy stanovich2008relative& 1 (numerical)& average number of purchases (numerical) in each condition (sunk cost yes vs. no)& 1, 4 \\ \midrule Allais problem stanovich2008relative&1 (binary choice) &lottery choice probability in each condition (form 1 vs. 2) &1, 4\\ \midrule Framing problem tversky1981framing& 1 (6-point Likert) &average preference for B vs. A (numerical) in each condition (framing gain vs. loss) &2, 4\\ \midrule Conjunction problem (Linda) tversky1983extensional& 3 (6-point Likert) & average prob. assessment in each condition (feminist bank teller vs. bank teller)& 2,4\\ \midrule Anchoring and adjustment tversky1974judgment,epley2004perspective& 2 (numerical) &average prediction (numerical) in each condition (with high vs. low anchor) &2, 4\\ \midrule Absolute vs. relative savings stanovich2008relative& 1 (binary choice) &probability of driving to store in each condition (calculator vs. jacket) &2,4\\ \midrule Myside bias stanovich2008relative& 1 (6-point Likert) &average ban agreement (numerical) in each condition (German vs. Ford) &2,4\\ \midrule Less is More stanovich2008relative& 3 (5/6-point Likert) &average attractiveness (numerical) in each condition (Form A vs. B vs. C) &3, 4\\ \midrule WTA/WTP – Thaler problem stanovich2008relative&1 (multiple choice)&average in each condition (WTP-certainty, WTA-certainty, WTP-noncertainty)& 3, 4 \\ \midrule \multicolumn{3}{l}{\emph{Heuristics and Biases (within subject)}} \\ \midrule False consensus furnas2024people& 10 (5-point Likert)+10 (slider) &average predicted public support for each level of own support &1,4 \\ \midrule Nonseparability of risk and benefits judgments stanovich2008relative& 8 (7-point Likert) & correlation between benefits and risks for each item &1,4 \\ \midrule Omission bias stanovich2008relative& 1 (4-point Likert) &likelihood of taking vaccine (numerical)& 2,4\\ \midrule Probability matching vs. maximizing stanovich2008relative& 6-10 (binary choice) &proportion choosing each strategy (Match, Max, other) &3, 4\\ \midrule Dominator neglect stanovich2008relative& 1 (binary choice) &proportion choosing large tray& 3,4\\ \midrule \multicolumn{3}{l}{\emph{Product Preferences}} \\ \midrule Pricing study gui2023challenge & 40 (binary choice) &demand curve for each product &3,4\\ \bottomrule \end{longtable}

Data

Table (ref) reports demographic characteristics of our sample. While our digital twins are created based on the raw responses, it is also informative and potentially useful to extract the measures corresponding to these questions (e.g., the extraversion score is measured by averaging 8 questions from the Big Five battery of questions). Across waves 1-3, we collected 47 measures capturing personality traits, cognitive abilities and economic preferences. Studying the correlations between these measures is also of interest to social scientists, above and beyond the question of digital twins. Technical appendix (ref) details the construction of these measures from the raw data, and Table (ref) reports summary statistics for the individual-level measures collected in the study. We compute a total of 1,326 pairwise correlations between the 47 measures listed in Table (ref) and 5 demographic characteristics. We apply the Bonferroni correction and consider a correlation as significant if the p-value is below $\frac{0.05}{1326}$. This gives us 509 pairs of measures with significant correlations. We cannot report them all here, and instead report in Table (ref) 10 examples of correlations that are particularly high and/or noteworthy. These correlations all have good face validity, which suggests the data are of high quality, despite the large number of questions.

Next, we test whether our 16 heuristics and biases experiments replicate known results at the aggregate level. See Table (ref). We see that both in waves 1-3 and in wave 4, all between-subject results replicate those in the literature, with the exception of the base rate fallacy. While kahneman1973psychology find that probability assessments are not sensitive to base rate, we find that they are. In terms of within-subject experiments, waves 1-3 and wave 4 also replicate all known results, with the exception of the non-separability of risks and benefits for one of the items, bicycles. While stanovich2008relative find a negative correlation between judged benefits and risks for this item, we find no correlation. The fact that our data replicates the vast majority of these known experimental results is again a sign of good data quality.

Finally, we calculate the test-retest accuracy in our data. We use the answers from waves 1-3 to our 88 holdout questions (across 177 tasks) as the ground truth. Given that all holdout questions are either binary or numerical (or transformed into numerical answers), we calculate accuracy as follows. For binary measures, accuracy is simply a binary indicator of whether two answers match. For non-binary measures, we calculate the absolute deviation between the ground truth and predicted answer, divided by the range of possible answers.\footnote{For the anchoring questions which accept unbounded answers, we transform the data into deciles based on the answers from wave 2 before calculating the absolute deviation.} We then compute accuracy as 1 minus this absolute deviation. This measure generalizes accuracy from binary to numerical questions: it ranges between 0 and 1, is equal to 1 when the prediction is equal to the ground truth, and 0 when it is maximally different. When multiple questions are included in the same task, we take the mean accuracy across the questions within each task. Therefore, we are left with one measure of accuracy per respondent for each of the 17 tasks (11 between-subject experiments, 5 within-subject experiments, 1 pricing study). Figure (ref) reports the average accuracy across respondents for each task as well as 95% confidence interval. We see that the average test-retest accuracy across the 17 tasks is 81.72%. This number is aligned with others reported in the literature (e.g., park2024generative), and again gives us confidence in the data's quality.

Creation of the digital twins

To construct each digital twin, we begin by merging the original Qualtrics survey files (QSF) with each participant’s raw responses, creating a self-contained JSON record for every individual. This record lists, in order, every question the participant actually encountered, the response options shown, and the answers. We then partition this record into three separate files:

itemize• Persona JSON: Aggregates all non–hold-out content from waves 1-3, used to define the persona. • Evaluation answer-block JSON: Contains the participant’s wave 1–3 responses to hold-out items, providing the ground truth for evaluating simulation accuracy. • Retest answer-block JSON: Stores wave 4 responses to those same hold-out items, used solely to compute the human test–retest accuracy benchmark.

By distributing the data in this modular format, we enable future researchers to experiment with alternative encoding or summarization schemes before presenting the material to an LLM.

In the present work, we use a straightforward text-based approach: the JSON files are converted into text descriptions detailing the questions, options, and participant answers. The prompt instruction is attached in Technical Appendix (ref). The model’s completion is then post-processed back into canonical survey coding and compared with the Wave 1–3 ground truth, yielding the accuracy statistics reported in Figure (ref). Our dataset includes all these files for future researchers to train and test LLM-based digital twin simulations.

Initial tests of digital twins' predictive performance

Individual level

To systematically evaluate different strategies for LLM-based persona simulation, we experimented with over a dozen variations in both persona construction and simulation methodology. These include differences in input format (e.g., text vs. JSON), model choices, prompting strategies such as proactive reasoning or chain-of-thought, persona summaries, and preliminary fine-tuning. Full experimental details are provided in Technical Appendix (ref). Overall, we find that the predictive accuracy of the answers simulated by the digital twins falls within a similar range across approaches (see Table (ref)). We hope this collection of baseline results will serve as a useful benchmark for future researchers exploring more advanced methods for persona training, such as reinforcement learning with human feedback (RLHF). For the initial analyses reported here, we focus on the text format using GPT4.1-mini.

table[table omitted — 982 chars of source]

Figure (ref) reports, for each task, the predictive accuracy of the answers simulated by the digital twins, as well as the accuracy of a random benchmark which chooses each answer from a random uniform distribution. On average, across the 17 tasks the accuracy of the digital twin predictions is 71.72%, and the ratio of the digital twin accuracy to the test-retest accuracy is 87.67%. The improvement over the baseline is consistent across all question types, highlighting the value of personalization and LLM-based simulation.

figure[figure omitted — 148 chars of source]

Aggregate level

We test whether the data simulated from the digital twins replicates the average treatment effects from the 11 classic between-subject studies and the 5 classic within-subject studies included in our experiment. Table (ref) shows that for 6 of the 10 results replicated by waves 1-3 and wave 4, the results from the digital twins also replicate the results. For anchoring and adjustment, the digital twins replicate the effect when asking participants to estimate the height of the highest redwood tree. But when asking participants to estimate the number of African countries in the UN, 98.8% of the twins gave the correct answer (54) and no anchoring effect was found. Three other between-subject effects were not replicated. In the outcome bias experiment, participants evaluate a physician's decision to operate on a patient. Humans evaluate the decision more favorably when the operation succeeded than when it failed, despite the risk being greater in the first condition. In contrast, digital twins all gave a favorable rating (“correct” to “clearly correct”), with no significant different across conditions. In the sunk cost fallacy experiment, the effect was actually reversed with the digital twins vs. their human counterparts, which we hope future research can explore. In the Allais problem experiment, which tests for violation of the independence axiom of utility theory, all digital twins chose the lower risk - lower reward option over the higher risk - higher reward one. Humans, on the other hand, were much more split in their decisions, and showed systematic differences across conditions (which violates the independence axiom of utility theory). Finally, the base rate fallacy, which was replicated neither in wave 2 nor in wave 4, was not replicated by the digital twins either.

Moving to within-subject experiments, we find that the digital twin results match the human results in two of the five within-subject experiments (see Table (ref)). In the nonseparability of risk and benefit judgments study, the digital twins judgments display negative correlation as predicted, but the correlation is significant for only one of the items. For probability matching vs. maximizing, the digital twins always selected the normative option while their human counterparts chose the normative option about 30% of the time only. For omission bias, participants were asked whether they would accept a vaccine that prevents catching a flu that has a 10% chance of killing affected patients when the vaccine itself carries a 5% chance of death. While approximately 45% of the human participants (45.10% in wave 2, 44.80% in wave 4) refused the vaccine, only 4.0% of the twins refused the vaccine. This finding echoes our finding related to outcome bias where digital twins were much more favorable to medical professionals compared to their human counterparts.

Finally, we construct average demand curves from the pricing study. Figure (ref) shows the average demand curves from the responses from waves 3 vs. wave 4 vs. digital twins. We see that the average demand curves from wave 3 vs. 4 are practically indistinguishable. We find that the average demand curve obtained from the twin is not fully downward sloping, due to the twins' responses to free products. This echoes gui2023challenge, although our digital twins produce demand curves that are downward sloping for positive prices and that are generally closer to the ground truth compared to the demand curves obtained by gui2023challenge without such input data.

In sum, while digital twins replicate most between-subject and within-subject effects, there are notable exceptions. Some occur when digital twins fail to mimic the suboptimal or non-normative behaviors of humans, or cannot “unlearn” certain facts (e.g., the number of African countries in the UN). In some areas, twins do match suboptimal human responses (e.g., absolute vs.\ relative savings, less is more, dominator neglect), raising the broader question of whether digital twins should be seen as “improved” humans or as models that also replicate human deviations from normative behavior and knowledge gaps.

Other deviations appear in the medical domain (e.g., outcome bias, omission bias). Future research should examine whether digital twins are systematically more trusting of the medical profession than humans. Another factor may be that certain topics, such as vaccination, have become highly polarized, making it difficult for GPT models to reflect conservative viewpoints motoki2024more. For instance, in our false consensus task, while about 45% of humans somewhat or strongly supported increased deportations, 74.1% of digital twins strongly or somewhat opposed the measure. More research is needed to systematically study where digital twins diverge from humans, especially in medical and political domains, and to identify other domains where such differences may arise.

\tiny

longtable[longtable omitted — 3,631 chars of source]

Conclusion

We present a unique dataset spanning over 500 questions and 2,000 respondents, with high data quality evidenced by strong correlations, good test-retest accuracy, and replication of known effects. While this resource can benefit social scientists broadly, our primary focus is on using it to build digital twins. These twins predict human behavior with out-of-sample accuracy reaching 87% of the test-retest benchmark. Replication of average treatment effects is generally good, though further research is needed to determine if digital twins can capture non-normative behaviors and reflect the full diversity of political and domain-specific views. The dataset’s focus on the US and social science topics is a potential limitation. Overall, we hope this resource accelerates LLM research and social science applications while being mindful of societal risks such as dehumanization of research and excessive reliance on AI in decision-making.