Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
39,782 characters · 8 sections · 91 citation commands
Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions
The rise of large language models (LLMs) like GPT has sparked interest across disciplines (including computer science, marketing, economics, psychology and political science) in leveraging these tools to create “silicon samples” which may replicate how these humans would behave in response to any stimuli aher_using_2023,argyle_out_2023, brand_using_2023,dillion_can_2023,hewitt2024predicting,horton_large_2023,li2025llm,park_generative_2023,qin2024aiturk. If these LLM-simulations can be a faithful substitute for eliciting responses from their human counterparts, the implications for both academics and practitioners are substantial. Academics could use silicon samples for pilot experiments to pinpoint stimuli with significant impact, thus improving the efficiency of theory development and experimental design. Firms could leverage these realistic simulations to explore different ideas and strategies, thereby improving customer insight and product development. Accordingly, in the recent past we have witnessed a large influx of firms offering services leveraging silicon samples for customer insights (e.g., \href{https://www.syntheticusers.com/}{Synthetic Users}, \href{https://outset.ai/}{Outset AI}, \href{https://www.nexxt.in/}{Nexxt}, \href{https://www.voxpopme.com/}{Voxpopme}, \href{https://www.evidenza.ai/}{Evidenza}, \href{https://www.expectedparrot.com/}{Expected Parrot}, \href{https://meaningful.app/}{Meaningful}, \href{https://xpolls.ai/}{xPolls}, \href{https://www.ipsos.com/en-uk/large-language-models-market-research}{Ipsos}, \href{https://civicsync.com/}{CivicSync}).
While silicon samples may be generated using only demographic information or hypothetical “life stories,” a promising approach consists in creating silicon samples that are “digital twins” of real people. Notably, park2024generative use LLMs to create digital twins of over 1,000 individuals based on transcripts from qualitative interviews, and find that the simulated agents replicated the human participants' responses on the General Social Survey 85% as accurately as participants replicate their own answers two weeks later.
Despite the promise and excitement surrounding digital twins, some uncertainty remains. For example, brucks2023prompt show that the answers provided by LLMs may be overly influenced by the architecture of the prompt, such as the labeling or ordering of options in multiple choice questions. gui2023challenge show that leveraging LLMs to simulate experiments may introduce unwanted confounding, due to the difficulty of clearly instructing the LLM how to draw variables not specified in the prompt. Other research santurkar_whose_2023,motoki2024more,li2025llm suggests LLMs tend to express opinions that are not representative of the (human) population.
Given this background, it is crucial for the academic and practitioner community to validate digital twins in a transparent, reliable, and replicable manner. However, existing datasets present significant limitations that hinder their effectiveness for this purpose. Some datasets are publicly available (e.g., alattar2018introduction,pew_research_2023,santurkar2023whose), but they are not well suited for testing the validity of digital twins because they do not contain behavioral data (e.g., experiments) nor a test-retest accuracy benchmark. Other datasets have this feature, but they are not publicly available (e.g., park2024generative).
In sum, there is no publicly available dataset that combines rich psychological profiles, behavioral data, and demographics from a large, representative sample. As a result, researchers often rely on synthetic or proprietary data, which undermines transparency, reliability, and replicability.
To address this gap, we assemble and publicly share an extensive dataset for a representative sample of $N=2,058$ people who each answered over 500 questions covering a wide range of demographic questions, psychological scales, cognitive performance questions, economic preferences questions, as well as replications of a wide range of within- and between-subject experiments on heuristics and biases taken from the behavioral economics literature. The data was collected across 4 waves of studies lasting on average 2.42 hour per participant in total. Table (ref) gives an overview of the measures collected in each wave and Figure (ref) illustrates our overall approach.
We use the responses to the heuristics and biases questions from waves 1-3 as holdout data, and train the digital twins based on the rest of the data from waves 1-3. Wave 4 repeated the heuristics and biases experiments, providing us with a measure of test-retest accuracy. Future uses of the data may keep the same split, or combine all the data from the 4 waves to create digital twins.
We report encouraging results regarding the quality of the data: correlations between measures have good face validity, we replicate almost all known results from the behavioral economics literature, and the test-retest accuracy is robust.
We also report initial tests of the predictive validity of digital twins constructed using the data. At the individual level, we compute the accuracy of the digital twin predictions on holdout questions, against the test-retest accuracy benchmark as well as a random benchmark. At the aggregate level, we test whether the digital twin simulations replicate the average treatment effects observed in human data. Throughout this process, we explore the type of behavior that can be predicted with higher vs. lower accuracy by the digital twins, to develop insight into the range of potential applications.
To the best of our knowledge, this dataset is unique in its scale and breadth. As such, there is value in the raw results, irrespective of the application to digital twins. We report descriptive statistics and correlations between the dozens of measures we collected. We encourage others to explore heterogeneous treatment effects dean2019empirical,stanovich2008relative.
We assembled a wide-range of measures proposed in the social science literature over the past several decades. In addition to 14 demographic questions, we included 19 personality tests that measured 26 constructs over 279 questions, 11 cognitive ability tests (85 questions, 11 measures), 10 economic preferences tests (34 questions, 10 measures). We also replicated 11 between-subject experiments (16 questions) and 5 within-subject experiments (32 questions) from the behavioral economics literature. Finally, we administered the pricing study from gui2023challenge, which asks participants to make purchase decisions about 40 different products at randomly selected prices. In total, participants answered 500 questions across the first three waves. Wave 4 repeated the within- and between-subject heuristics and biases experiments from the first three waves as well as the pricing study from wave 3 (88 questions in total). Participants were assigned to the exact same condition in wave 4 as they were in waves 1-3 for each of these experiments, providing us with a clean measure of test-retest accuracy. We programmed the studies on Qualtrics, doing our best to replicate the stimuli and measures from the original papers.\footnote{We made minor adjustments to reflect cultural and societal changes (e.g., in the mental accounting scenarios from thaler1985mental we replaced “Mr. A” and “Mr. B” with “Person A” and “Person B,” and in the sunk cost experiment of stanovich2008relative we replaced video rental stores with coffee shops).} Our study was approved as IRB Protocol (blinded for anonymity).\footnote{Our study was determined as exempt on the basis of “involving benign behavioral interventions in conjunction with the collection of information from an adult subject through verbal or written responses.”}
We launched Wave 1 on Prolific on 01/29/2025, targeting 2,500 representative US respondents (sampled by age, sex, and ethnicity). Participants received \$7 for completing Wave 1.\footnote{We pre-tested each wave to estimate response time and adjusted compensation accordingly.} They were informed that this was the first of four waves and would earn a \$10 bonus for completing all waves (with comprehension checks to ensure understanding). We received 2,509 complete responses. The following week (02/04/2025), we invited these 2,509 participants to Wave 2 (\$7), receiving 2,263 responses. The next week (02/11/2025), we invited them again for Wave 3 (\$7), receiving 2,252 responses. Waves 2 and 3 were closed the next week. On 02/25/2025, we invited the 2,154 participants who had completed Waves 1–3 to Wave 4 (\$6; two-week delay since Wave 4 repeated previous measures), receiving 2,058 responses and closing the wave after one week. Those completing all four waves received an additional \$10 bonus, totaling \$37.
These 2,058 participants who completed all four waves constitute our final sample. Among our final sample, the average response time was 43.88 minutes for wave 1 (std=19.27), 45.31 minutes for wave 2 (std=19.24), 32.66 minutes for wave 3 (std=15.68), and 24.09 minutes for wave 4 (std=12.51). The average total time across all 4 waves was 145.47 minutes (std=56.10).
Table (ref) reports demographic characteristics of our sample. While our digital twins are created based on the raw responses, it is also informative and potentially useful to extract the measures corresponding to these questions (e.g., the extraversion score is measured by averaging 8 questions from the Big Five battery of questions). Across waves 1-3, we collected 47 measures capturing personality traits, cognitive abilities and economic preferences. Studying the correlations between these measures is also of interest to social scientists, above and beyond the question of digital twins. Technical appendix (ref) details the construction of these measures from the raw data, and Table (ref) reports summary statistics for the individual-level measures collected in the study. We compute a total of 1,326 pairwise correlations between the 47 measures listed in Table (ref) and 5 demographic characteristics. We apply the Bonferroni correction and consider a correlation as significant if the p-value is below $\frac{0.05}{1326}$. This gives us 509 pairs of measures with significant correlations. We cannot report them all here, and instead report in Table (ref) 10 examples of correlations that are particularly high and/or noteworthy. These correlations all have good face validity, which suggests the data are of high quality, despite the large number of questions.
Next, we test whether our 16 heuristics and biases experiments replicate known results at the aggregate level. See Table (ref). We see that both in waves 1-3 and in wave 4, all between-subject results replicate those in the literature, with the exception of the base rate fallacy. While kahneman1973psychology find that probability assessments are not sensitive to base rate, we find that they are. In terms of within-subject experiments, waves 1-3 and wave 4 also replicate all known results, with the exception of the non-separability of risks and benefits for one of the items, bicycles. While stanovich2008relative find a negative correlation between judged benefits and risks for this item, we find no correlation. The fact that our data replicates the vast majority of these known experimental results is again a sign of good data quality.
Finally, we calculate the test-retest accuracy in our data. We use the answers from waves 1-3 to our 88 holdout questions (across 177 tasks) as the ground truth. Given that all holdout questions are either binary or numerical (or transformed into numerical answers), we calculate accuracy as follows. For binary measures, accuracy is simply a binary indicator of whether two answers match. For non-binary measures, we calculate the absolute deviation between the ground truth and predicted answer, divided by the range of possible answers.\footnote{For the anchoring questions which accept unbounded answers, we transform the data into deciles based on the answers from wave 2 before calculating the absolute deviation.} We then compute accuracy as 1 minus this absolute deviation. This measure generalizes accuracy from binary to numerical questions: it ranges between 0 and 1, is equal to 1 when the prediction is equal to the ground truth, and 0 when it is maximally different. When multiple questions are included in the same task, we take the mean accuracy across the questions within each task. Therefore, we are left with one measure of accuracy per respondent for each of the 17 tasks (11 between-subject experiments, 5 within-subject experiments, 1 pricing study). Figure (ref) reports the average accuracy across respondents for each task as well as 95% confidence interval. We see that the average test-retest accuracy across the 17 tasks is 81.72%. This number is aligned with others reported in the literature (e.g., park2024generative), and again gives us confidence in the data's quality.
To construct each digital twin, we begin by merging the original Qualtrics survey files (QSF) with each participant’s raw responses, creating a self-contained JSON record for every individual. This record lists, in order, every question the participant actually encountered, the response options shown, and the answers. We then partition this record into three separate files:
By distributing the data in this modular format, we enable future researchers to experiment with alternative encoding or summarization schemes before presenting the material to an LLM.
In the present work, we use a straightforward text-based approach: the JSON files are converted into text descriptions detailing the questions, options, and participant answers. The prompt instruction is attached in Technical Appendix (ref). The model’s completion is then post-processed back into canonical survey coding and compared with the Wave 1–3 ground truth, yielding the accuracy statistics reported in Figure (ref). Our dataset includes all these files for future researchers to train and test LLM-based digital twin simulations.
To systematically evaluate different strategies for LLM-based persona simulation, we experimented with over a dozen variations in both persona construction and simulation methodology. These include differences in input format (e.g., text vs. JSON), model choices, prompting strategies such as proactive reasoning or chain-of-thought, persona summaries, and preliminary fine-tuning. Full experimental details are provided in Technical Appendix (ref). Overall, we find that the predictive accuracy of the answers simulated by the digital twins falls within a similar range across approaches (see Table (ref)). We hope this collection of baseline results will serve as a useful benchmark for future researchers exploring more advanced methods for persona training, such as reinforcement learning with human feedback (RLHF). For the initial analyses reported here, we focus on the text format using GPT4.1-mini.
Figure (ref) reports, for each task, the predictive accuracy of the answers simulated by the digital twins, as well as the accuracy of a random benchmark which chooses each answer from a random uniform distribution. On average, across the 17 tasks the accuracy of the digital twin predictions is 71.72%, and the ratio of the digital twin accuracy to the test-retest accuracy is 87.67%. The improvement over the baseline is consistent across all question types, highlighting the value of personalization and LLM-based simulation.
We test whether the data simulated from the digital twins replicates the average treatment effects from the 11 classic between-subject studies and the 5 classic within-subject studies included in our experiment. Table (ref) shows that for 6 of the 10 results replicated by waves 1-3 and wave 4, the results from the digital twins also replicate the results. For anchoring and adjustment, the digital twins replicate the effect when asking participants to estimate the height of the highest redwood tree. But when asking participants to estimate the number of African countries in the UN, 98.8% of the twins gave the correct answer (54) and no anchoring effect was found. Three other between-subject effects were not replicated. In the outcome bias experiment, participants evaluate a physician's decision to operate on a patient. Humans evaluate the decision more favorably when the operation succeeded than when it failed, despite the risk being greater in the first condition. In contrast, digital twins all gave a favorable rating (“correct” to “clearly correct”), with no significant different across conditions. In the sunk cost fallacy experiment, the effect was actually reversed with the digital twins vs. their human counterparts, which we hope future research can explore. In the Allais problem experiment, which tests for violation of the independence axiom of utility theory, all digital twins chose the lower risk - lower reward option over the higher risk - higher reward one. Humans, on the other hand, were much more split in their decisions, and showed systematic differences across conditions (which violates the independence axiom of utility theory). Finally, the base rate fallacy, which was replicated neither in wave 2 nor in wave 4, was not replicated by the digital twins either.
Moving to within-subject experiments, we find that the digital twin results match the human results in two of the five within-subject experiments (see Table (ref)). In the nonseparability of risk and benefit judgments study, the digital twins judgments display negative correlation as predicted, but the correlation is significant for only one of the items. For probability matching vs. maximizing, the digital twins always selected the normative option while their human counterparts chose the normative option about 30% of the time only. For omission bias, participants were asked whether they would accept a vaccine that prevents catching a flu that has a 10% chance of killing affected patients when the vaccine itself carries a 5% chance of death. While approximately 45% of the human participants (45.10% in wave 2, 44.80% in wave 4) refused the vaccine, only 4.0% of the twins refused the vaccine. This finding echoes our finding related to outcome bias where digital twins were much more favorable to medical professionals compared to their human counterparts.
Finally, we construct average demand curves from the pricing study. Figure (ref) shows the average demand curves from the responses from waves 3 vs. wave 4 vs. digital twins. We see that the average demand curves from wave 3 vs. 4 are practically indistinguishable. We find that the average demand curve obtained from the twin is not fully downward sloping, due to the twins' responses to free products. This echoes gui2023challenge, although our digital twins produce demand curves that are downward sloping for positive prices and that are generally closer to the ground truth compared to the demand curves obtained by gui2023challenge without such input data.
In sum, while digital twins replicate most between-subject and within-subject effects, there are notable exceptions. Some occur when digital twins fail to mimic the suboptimal or non-normative behaviors of humans, or cannot “unlearn” certain facts (e.g., the number of African countries in the UN). In some areas, twins do match suboptimal human responses (e.g., absolute vs.\ relative savings, less is more, dominator neglect), raising the broader question of whether digital twins should be seen as “improved” humans or as models that also replicate human deviations from normative behavior and knowledge gaps.
Other deviations appear in the medical domain (e.g., outcome bias, omission bias). Future research should examine whether digital twins are systematically more trusting of the medical profession than humans. Another factor may be that certain topics, such as vaccination, have become highly polarized, making it difficult for GPT models to reflect conservative viewpoints motoki2024more. For instance, in our false consensus task, while about 45% of humans somewhat or strongly supported increased deportations, 74.1% of digital twins strongly or somewhat opposed the measure. More research is needed to systematically study where digital twins diverge from humans, especially in medical and political domains, and to identify other domains where such differences may arise.
\tiny
We present a unique dataset spanning over 500 questions and 2,000 respondents, with high data quality evidenced by strong correlations, good test-retest accuracy, and replication of known effects. While this resource can benefit social scientists broadly, our primary focus is on using it to build digital twins. These twins predict human behavior with out-of-sample accuracy reaching 87% of the test-retest benchmark. Replication of average treatment effects is generally good, though further research is needed to determine if digital twins can capture non-normative behaviors and reflect the full diversity of political and domain-specific views. The dataset’s focus on the US and social science topics is a potential limitation. Overall, we hope this resource accelerates LLM research and social science applications while being mindful of societal risks such as dehumanization of research and excessive reliance on AI in decision-making.