EconBase
← Back to paper

Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

39,782 characters

Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions




\maketitle


\begin{abstract}
LLM-based digital twin simulation, where large language models are used to emulate individual human behavior, holds great promise for research in AI, social science, and digital experimentation. However, progress in this area has been hindered by the scarcity of real, individual-level datasets that are both large and publicly available. This lack of high-quality ground truth limits both the development and validation of digital twin methodologies. To address this gap, we introduce a large-scale, public dataset designed to capture a rich and holistic view of individual human behavior. We survey a representative sample of $N = 2,058$ participants (average 2.42 hours per person) in the US across four waves with 500 questions in total, covering a comprehensive battery of demographic, psychological, economic, personality, and cognitive measures, as well as replications of behavioral economics experiments and a pricing survey. The final wave repeats tasks from earlier waves to establish a test-retest accuracy baseline. Initial analyses suggest the data are of high quality and show promise for constructing digital twins that predict human behavior well at the individual and aggregate levels. By making the full dataset publicly available, we aim to establish a valuable testbed for the development and benchmarking of LLM-based persona simulations. Beyond LLM applications, due to its unique breadth and scale the dataset also enables broad social science research, including studies of cross-construct correlations and heterogeneous treatment effects.~\footnote{The dataset is publicly available at \url{https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500/}, and the LLM simulation code can be found at \url{https://github.com/tianyipeng-lab/Digital-Twin-Simulation}.
}
\end{abstract}

\section{Introduction}
The rise of large language models (LLMs) like GPT has sparked interest across disciplines (including computer science, marketing,  economics, psychology and political science) in leveraging these tools to create ``silicon samples'' which may replicate how these humans would behave in response to any stimuli \citep{aher_using_2023,argyle_out_2023, brand_using_2023,dillion_can_2023,hewitt2024predicting,horton_large_2023,li2025llm,park_generative_2023,qin2024aiturk}.
If these LLM-simulations can be a faithful substitute for eliciting responses from their human counterparts, the implications for both academics and practitioners are substantial. Academics could use silicon samples for pilot experiments to pinpoint stimuli with significant impact, thus improving the efficiency of theory development and experimental design.
Firms could leverage these realistic simulations to explore different ideas and strategies, thereby improving customer insight and product development. Accordingly, in the recent past we have witnessed a large influx of firms offering services leveraging silicon samples for customer insights (e.g., \href{https://www.syntheticusers.com/}{Synthetic Users}, \href{https://outset.ai/}{Outset AI}, \href{https://www.nexxt.in/}{Nexxt}, \href{https://www.voxpopme.com/}{Voxpopme}, \href{https://www.evidenza.ai/}{Evidenza}, \href{https://www.expectedparrot.com/}{Expected Parrot}, \href{https://meaningful.app/}{Meaningful}, \href{https://xpolls.ai/}{xPolls}, \href{https://www.ipsos.com/en-uk/large-language-models-market-research}{Ipsos}, \href{https://civicsync.com/}{CivicSync}).

While silicon samples may be generated using only demographic information or hypothetical ``life stories,'' a promising approach consists in creating silicon samples that are ``digital twins'' of real people. Notably, \cite{park2024generative}
use LLMs to create digital twins of over 1,000 individuals based on transcripts from qualitative interviews, and find that the simulated agents replicated the human participants' responses on the General Social Survey 85\% as accurately as participants replicate their own answers two weeks later.

Despite the promise and excitement surrounding digital twins, some uncertainty remains. For example, \cite{brucks2023prompt} show that the answers provided by LLMs may be overly influenced by the architecture of the prompt, such as the labeling or ordering of options in multiple choice questions. \cite{gui2023challenge} show that leveraging LLMs to simulate experiments may introduce unwanted confounding, due to the difficulty of clearly instructing the LLM how to draw variables not specified in the prompt. Other research  \citep{santurkar_whose_2023,motoki2024more,li2025llm} suggests LLMs tend to express opinions that are not representative of the (human) population.

Given this background, it is crucial for the academic and practitioner community to validate digital twins in a transparent, reliable, and replicable manner. However, existing datasets present significant limitations that hinder their effectiveness for this purpose. Some datasets are publicly available (e.g., \citealp{alattar2018introduction,pew_research_2023,santurkar2023whose}), but they are not well suited for testing the validity of digital twins because they do not contain behavioral data (e.g., experiments) nor a test-retest accuracy benchmark. Other datasets have this feature, but they are not publicly available (e.g., \citealp{park2024generative}).


In sum, there is no publicly available dataset that combines rich psychological profiles, behavioral data, and demographics from a large, representative sample. As a result, researchers often rely on synthetic or proprietary data, which undermines transparency, reliability, and replicability.



To address this gap, we assemble and publicly share an extensive dataset for a representative sample of $N=2,058$ people who each answered over 500 questions covering a wide range of demographic questions, psychological scales, cognitive performance questions, economic preferences questions, as well as replications of a wide range of within- and between-subject experiments on heuristics and biases taken from the behavioral economics literature. The data was collected across 4 waves of studies lasting on average
2.42 hour per participant in total. Table \ref{tab:measures} gives an overview of the measures collected in each wave and Figure \ref{fig:overview} illustrates our overall approach.





We use the responses to the heuristics and biases questions from waves 1-3 as holdout data, and train the digital twins based on the rest of the data from waves 1-3. Wave 4 repeated the heuristics and biases experiments, providing us with a measure of test-retest accuracy. Future uses of the data may keep the same split, or combine all the data from the 4 waves to create digital twins.

We report encouraging results regarding the quality of the data: correlations between measures have good face validity, we replicate almost all known results from the behavioral economics literature, and the test-retest accuracy is robust.

We also report initial tests of the predictive validity of digital twins constructed using the data.
At the individual level, we compute the accuracy of the digital twin predictions on holdout questions, against the test-retest accuracy benchmark as well as a random benchmark.
At the aggregate level, we test whether the digital twin simulations replicate the average treatment effects observed in human data. Throughout this process, we explore the type of behavior that can be predicted with higher vs. lower accuracy by the digital twins, to develop insight into the range of potential applications.

To the best of our knowledge, this dataset is unique in its scale and breadth. As such, there is value in the raw results, irrespective of the application to digital twins. We report descriptive statistics and correlations between the dozens of measures we collected. We encourage others to explore heterogeneous treatment effects \citep{dean2019empirical,stanovich2008relative}.



\begin{figure}[htbp]
  \centering
  \includegraphics[width=\textwidth,trim={5cm 9cm 5cm 3cm},clip]{figs/Digital_Twins_survey_v2.pdf}
  \caption{Overview}
  \label{fig:overview}
\end{figure}


\section{Methods}

We assembled a wide-range of measures proposed in the social science literature over the past several decades. In addition to 14 demographic questions, we included 19 personality tests that measured 26 constructs over 279 questions, 11 cognitive ability tests (85 questions, 11 measures), 10 economic preferences tests (34 questions, 10 measures). We also replicated 11 between-subject experiments (16 questions) and 5 within-subject experiments (32 questions) from the behavioral economics literature. Finally, we administered the pricing study from \cite{gui2023challenge}, which asks participants to make purchase decisions about 40 different products at randomly selected prices. In total, participants answered 500 questions across the first three waves. Wave 4 repeated the within- and between-subject heuristics and biases experiments from the first three waves as well as the pricing study from wave 3 (88 questions in total). Participants were assigned to the exact same condition in wave 4 as they were in waves 1-3 for each of these experiments, providing us with a clean measure of test-retest accuracy.
We programmed the studies on Qualtrics, doing our best to replicate the stimuli and measures from the original papers.\footnote{We made minor adjustments to reflect cultural and societal changes (e.g., in the mental accounting scenarios from \cite{thaler1985mental} we replaced ``Mr. A'' and ``Mr. B'' with ``Person A'' and ``Person B,'' and in the sunk cost experiment of \cite{stanovich2008relative} we replaced video rental stores with coffee shops).} Our study was approved as IRB Protocol (blinded for anonymity).\footnote{Our study was determined as exempt on the basis of ``involving benign behavioral interventions in conjunction with the collection of information from an adult subject through verbal or written responses.''}

We launched Wave~1 on Prolific on 01/29/2025, targeting 2,500 representative US respondents (sampled by age, sex, and ethnicity). Participants received \$7 for completing Wave~1.\footnote{We pre-tested each wave to estimate response time and adjusted compensation accordingly.} They were informed that this was the first of four waves and would earn a \$10 bonus for completing all waves (with comprehension checks to ensure understanding). We received 2,509 complete responses. The following week (02/04/2025), we invited these 2,509 participants to Wave~2 (\$7), receiving 2,263 responses. The next week (02/11/2025), we invited them again for Wave~3 (\$7), receiving 2,252 responses. Waves~2 and~3 were closed the next week. On 02/25/2025, we invited the 2,154 participants who had completed Waves~1–3 to Wave~4 (\$6; two-week delay since Wave~4 repeated previous measures), receiving 2,058 responses and closing the wave after one week. Those completing all four waves received an additional \$10 bonus, totaling \$37.



These 2,058 participants who completed all four waves constitute our final sample.  Among our final sample, the average response time was 43.88 minutes for wave 1 (std=19.27), 45.31 minutes for wave 2 (std=19.24), 32.66 minutes for wave 3 (std=15.68), and 24.09 minutes for wave 4 (std=12.51). The average total time across all 4 waves was 145.47 minutes (std=56.10).

\begin{landscape}
    \centering
    \tiny
    \begin{longtable}{p{6.5cm}p{2.5cm}p{8cm}p{1cm}}
\caption{Complete list of questions and related measures}
\label{tab:measures}\\
\toprule

                \textbf{Task (source)} & \textbf{\#Questions (format)} & \textbf{Extracted measure(s)} & \textbf{Wave(s)} \\
                \midrule
                        \multicolumn{4}{l}{\emph{Demographics}} \\
        \midrule
        Demographics \citep{santurkar_whose_2023}& 12 (multiple choice) & region, sex, age, education, race, citizenship, marital status, religion, religious attendance, political party, household income, political ideology (categorical) & 1 \\
        \midrule
        Additional demographics & 2 (multiple choice) & household size, employment status (categorical) & 1 \\
        \midrule

        \multicolumn{3}{l}{\emph{Personality Traits}} \\
        \midrule
       Big 5 personality test \citep{john1999big} & 44 (5-point Likert)  & extraversion, agreeableness, conscientiousness, neuroticism, openness scores (numerical) & 1 \\
       \midrule
        Need for cognition scale \citep{cacioppo1984efficient} & 18 (5-point Likert)  & need for cognition score (numerical) & 1 \\
        \midrule
        Agentic vs. Communal Values scale \citep{trapnell2012agentic} & 24 (9-point Likert)  & agency score, communion score (numerical) & 1 \\
        \midrule
        Consumer Minimalism scale \citep{wilson2022consumer} & 12 (5-point Likert)  & minimalism score (numerical) & 1 \\
        \midrule
        Empathy scale \citep{carre2013basic} & 20 (5-point Likert) & basic empathy score (numerical) & 1 \\
        \midrule
        Green values scale \citep{haws2014seeing} & 6 (5-point Likert) & green score (numerical) & 1 \\
        \midrule
        Social Desirability scale \citep{reynolds1982development} & 13 (binary choice) & social desirability score (numerical) & 2 \\
        \midrule
        Conscientiousness scale \citep{johnson2019s} & 8 (9-point Likert) & conscientiousness score (wave 2) (numerical) & 2 \\
        \midrule
        Anxiety scale \citep{beck1988inventory} & 21 (4-point Likert) & anxiety score (numerical) & 2 \\
        \midrule
        Individualism vs. Collectivism scale  \citep{triandis1998converging} & 16 (5-point Likert) & horizontal/vertical individualism, horizontal/vertical collectivism scores (numerical) & 2 \\
\midrule
        Selves questionnaire \citep{higgins1985self}&	3  (open-ended)&	n/a	& 2 \\
        \midrule
        Regulatory Focus scale \citep{fellner2007regulatory}&	10 (7-point Likert) &regulatory focus score (numerical)&3 \\
        \midrule
        Tightwads vs. Spendthrift scale	 \citep{rick2008tightwads}&	4 (multiple choice) &tightwads vs. spendthrift score (numerical)&3 \\
        \midrule
        Depression scale \citep{date1987beck}&22 (multiple choice)&depression score (numerical)&3 \\
        \midrule
        Need for uniqueness	scale \citep{ruvio2008consumers}&	12	(5-point Likert) &need for uniqueness score (numerical)&3 \\
        \midrule
        Self-monitoring	scale \citep{lennox1984revision}&13 (6-point Likert)&self-monitoring score (numerical)	&3 \\
        \midrule
        Self-concept clarity scale \citep{campbell1996self}&12	(5-point Likert)&	self-concept clarity score (numerical)&	3 \\
        \midrule
        Need for closure scale	\citep{roets2011item}&	15 (5-point Likert) &need for closure score (numerical)&3 \\
        \midrule
        Maximization scale	\citep{nenkov2008short} &	6 (5-point Likert) 	&maximization score (numerical)	&3\\

        \midrule
 \multicolumn{3}{l}{\emph{Cognitive Abilities}} \\
        \midrule
        Cognitive Reflection Test \citep{krefeld2024exposing}& 4 (open-ended) & CRT score (numerical) & 1 \\
        \midrule
        Fluid intelligence test \citep{krefeld2024exposing} & 6 (multiple choice)      & fluid intelligence score (numerical) & 1 \\
        \midrule
        Crystallized intelligence test \citep{krefeld2024exposing} & 20 (multiple choice)  & crystallized intelligence score (numerical) & 1 \\
        \midrule
        Syllogisms test \citep{markovits1989belief} & 12 (multiple choice) & syllogism score (numerical) & 1 \\
        \midrule
        Overconfidence \citep{dean2019empirical}&	1 (numerical) &overconfidence score (own predicted-actual score)&	1\\
        \midrule
        Overplacement \citep{dean2019empirical}&	1 (numerical) &	overplacement score (own predicted score-predicted average)&	1\\
        \midrule
        Financial literacy test	\citep{johnson2019s}&	7 (mult. choice)+1 (num.)&financial literacy score (numerical)&2 \\
        \midrule
        Numeracy test \citep{johnson2019s}&	8 (numerical)&numeracy score (numerical)	&2 \\
        \midrule
        Deductive certainty of Modus Ponens test \citep{stanovich2008relative}&	4 (binary choice) 	&deductive certainty score&	2\\
        \midrule
        Forward Flow (free associations) \citep{gray2019forward}&	20 (open-ended)	&forward flow score (average pairwise semantic distance)&	2\\
        \midrule
        Wason Selection Task \citep{klauer2007abstract}&	1 (multiple choice) &Wason Selection Task score (numerical)	&3 \\
        \midrule

 \multicolumn{3}{l}{\emph{Economic Preferences}} \\
        \midrule
        Ultimatum game (sender) \citep{guth1982experimental} & 1 (multiple choice)  & ultimatum-send (percentage sent) & 1 \\
        \midrule
        Ultimatum game (receiver) \citep{guth1982experimental} & 6 (binary choice)  & ultimatum-receive (acceptance probability) & 1 \\
        \midrule
        Mental accounting \citep{thaler1985mental}& 4 (binary choice)  & mental accounting score (\% choices consistent with mental account predictions) & 1 \\
     \midrule
        Discount \citep{dean2019empirical} & 3 (multiple price list) & discount rate (numerical) & 2 \\
        \midrule
        Present bias \citep{dean2019empirical} & 3 (multiple price list) & present bias (numerical) & 2 \\
        \midrule
        Risk Aversion \citep{dean2019empirical} & 3 (uncertainty equivalence) & risk aversion coefficient (numerical) & 2 \\
        \midrule
        Loss Aversion \citep{dean2019empirical} & 4 (uncertainty equivalence) & loss aversion coefficient (numerical) & 2 \\
        \midrule
        Trust game (sender) \citep{dean2019empirical} & 1 (multiple choice) & trust-send (percentage sent) & 2 \\
         \midrule
        Trust game (receiver) \citep{dean2019empirical} & 5 (multiple choice) & Trust-return (average percentage returned) & 2 \\
        \midrule
        Trust game (sender) thought listing	&	1 (open-ended)	&n/a&	2\\
        \midrule
        Trust game (receiver) thought listing	&	1 (open-ended)	&n/a&	2\\
        \midrule
        Dictator game \citep{baron1988outcome}&	1 (multiple choice) &	dictator-send (percentage sent)&	3\\
        \midrule
        Dictator game thought listing		&1	(open-ended)	&n/a	&3\\
        \midrule

 \multicolumn{3}{l}{\emph{Heuristics and Biases (between subject)}} \\
        \midrule
         Base rate problem \citep{kahneman1973psychology}&	1 (slider scale) &average prob. assessment (numerical) in each condition (base rate of 30 vs. 70 engineers)	&1, 4\\
         \midrule
         Outcome bias \citep{baron1988outcome}&1 (7-point Likert) 	&average correctness assessment (numerical) in each condition (success vs. failure)&	1, 4 \\
         \midrule
        Sunk cost fallacy \citep{stanovich2008relative}&	1 (numerical)&	average number of purchases (numerical) in each condition (sunk cost yes vs. no)&	1, 4 \\
        \midrule
        Allais problem \citep{stanovich2008relative}&1	(binary choice)	&lottery choice probability in each condition (form 1 vs. 2)	&1, 4\\
        \midrule
        Framing problem	\citep{tversky1981framing}&	1 (6-point Likert) &average preference for B vs. A (numerical) in each condition (framing gain vs. loss)	&2, 4\\
        \midrule
        Conjunction problem (Linda) \citep{tversky1983extensional}&	3 (6-point Likert) &	average prob. assessment in each condition (feminist bank teller vs. bank teller)&	2,4\\
        \midrule
        Anchoring and adjustment \citep{tversky1974judgment,epley2004perspective}&	2 (numerical)	&average prediction (numerical) in each condition (with high vs. low anchor)	&2, 4\\
        \midrule
        Absolute vs. relative savings	\citep{stanovich2008relative}&	1 (binary choice)	&probability of driving to store in each condition (calculator vs. jacket)	&2,4\\
        \midrule
        Myside bias	\citep{stanovich2008relative}&	1	(6-point Likert) 	&average ban agreement (numerical) in each condition (German vs. Ford)	&2,4\\
        \midrule
        Less is More \citep{stanovich2008relative}&	3 (5/6-point Likert) &average attractiveness (numerical) in each condition (Form A vs. B vs. C)	&3, 4\\
        \midrule
        WTA/WTP – Thaler problem \citep{stanovich2008relative}&1 (multiple choice)&average  in each condition (WTP-certainty, WTA-certainty, WTP-noncertainty)&	3, 4 \\
        \midrule
        \multicolumn{3}{l}{\emph{Heuristics and Biases (within subject)}} \\
        \midrule
        False consensus	\citep{furnas2024people}&	10 (5-point Likert)+10 (slider)	&average predicted public support for each level of own support	&1,4 \\
        \midrule
        Nonseparability of risk and benefits judgments \citep{stanovich2008relative}&	8 (7-point Likert) &	correlation between benefits and risks for each item	&1,4 \\
        \midrule
        Omission bias \citep{stanovich2008relative}&	1 (4-point Likert) &likelihood of taking vaccine (numerical)&	2,4\\
        \midrule
        Probability matching vs. maximizing	\citep{stanovich2008relative}&	6-10 (binary choice)	&proportion choosing each strategy (Match, Max, other) &3, 4\\
        \midrule
        Dominator neglect \citep{stanovich2008relative}&	1	(binary choice)	&proportion choosing large tray&	3,4\\


        \midrule
        \multicolumn{3}{l}{\emph{Product Preferences}} \\
        \midrule
        Pricing study \citep{gui2023challenge}	&	40 (binary choice)	&demand curve for each product	&3,4\\
        \bottomrule

\end{longtable}
\end{landscape}

\section{Data}

Table \ref{tab:demo} reports demographic characteristics of our sample. While our digital twins are created based on the raw responses, it is also informative and potentially useful to extract the measures corresponding to these questions (e.g., the extraversion score is measured by averaging 8 questions from the Big Five battery of questions). Across waves 1-3, we collected 47 measures capturing personality traits, cognitive abilities and economic preferences.
Studying the correlations between these measures is also of interest to social scientists, above and beyond the question of digital twins. Technical appendix \ref{measures} details the construction of these measures from the raw data, and Table \ref{tab:summary} reports summary statistics for the individual-level measures collected in the study. We compute a total of 1,326 pairwise correlations between the 47 measures listed in Table \ref{tab:measures} and 5 demographic characteristics. We apply the Bonferroni correction and consider a correlation as significant if the \textit{p}-value is below $\frac{0.05}{1326}$. This gives us 509 pairs of measures with significant correlations. We cannot report them all here, and instead report in Table \ref{tab:correlations} 10 examples of correlations that are particularly high and/or noteworthy. These correlations all have good face validity, which suggests the data are of high quality, despite the large number of questions.

Next, we test whether our 16 heuristics and biases experiments replicate known results at the aggregate level. See Table \ref{tab:replication}. We see that both in waves 1-3 and in wave 4, all between-subject results replicate those in the literature, with the exception of the base rate fallacy. While \cite{kahneman1973psychology} find that probability assessments are not sensitive to base rate, we find that they are. In terms of within-subject experiments, waves 1-3 and wave 4 also replicate all known results, with the exception of the non-separability of risks and benefits for one of the items, bicycles. While \cite{stanovich2008relative} find a negative correlation between judged benefits and risks for this item, we find no correlation. The fact that our data replicates the vast majority of these known experimental results is again a sign of good data quality.

Finally, we calculate the test-retest accuracy in our data. We use the answers from waves 1-3 to our 88 holdout questions (across 177 tasks) as the ground truth. Given that all holdout questions are either binary or numerical (or transformed into numerical answers), we calculate accuracy as follows.
For binary measures, accuracy is simply a binary indicator of whether two answers match. For non-binary measures, we calculate the absolute deviation between the ground truth and predicted answer, divided by the range of possible answers.\footnote{For the anchoring questions which accept unbounded answers, we transform the data into deciles based on the answers from wave 2 before calculating the absolute deviation.} We then compute accuracy as 1 minus this absolute deviation. This measure generalizes accuracy from binary to numerical questions: it ranges between 0 and 1, is equal to 1 when the prediction is equal to the ground truth, and 0 when it is maximally different. When multiple questions are included in the same task, we take the mean accuracy across the questions within each task. Therefore, we are left with one measure of accuracy per respondent for each of the 17 tasks (11 between-subject experiments, 5 within-subject experiments, 1 pricing study). Figure \ref{fig:accuracy} reports the average accuracy across respondents for each task as well as 95\% confidence interval. We see that the average test-retest accuracy across the 17 tasks is 81.72\%. This number is aligned with others reported in the literature (e.g., \citealp{park2024generative}), and again gives us confidence in the data's quality.


\section{Creation of the digital twins}
To construct each digital twin, we begin by merging the original Qualtrics survey files (QSF) with each participant’s raw responses, creating a self-contained JSON record for every individual. This record lists, in order, every question the participant actually encountered, the response options shown, and the answers. We then partition this record into three separate files:
\begin{itemize}
    \item \textbf{Persona JSON:} Aggregates all non–hold-out content from waves 1-3, used to define the persona.
    \item \textbf{Evaluation answer-block JSON:} Contains the participant’s wave 1–3 responses to hold-out items, providing the ground truth for evaluating simulation accuracy.
    \item \textbf{Retest answer-block JSON:} Stores wave 4 responses to those same hold-out items, used solely to compute the human test–retest accuracy benchmark.
\end{itemize}
By distributing the data in this modular format, we enable future researchers to experiment with alternative encoding or summarization schemes before presenting the material to an LLM.

In the present work, we use a straightforward text-based approach: the JSON files are converted into text descriptions detailing the questions, options, and participant answers. The prompt instruction is attached in Technical Appendix \ref{sec:approach-details}. The model’s completion is then post-processed back into canonical survey coding and compared with the Wave 1–3 ground truth, yielding the accuracy statistics reported in Figure~\ref{fig:accuracy}. Our dataset includes all these files for future researchers to train and test LLM-based digital twin simulations.


\section{Initial tests of digital twins' predictive performance}

\subsection{Individual level}

To systematically evaluate different strategies for LLM-based persona simulation, we experimented with over a dozen variations in both persona construction and simulation methodology. These include differences in input format (e.g., text vs. JSON), model choices, prompting strategies such as proactive reasoning or chain-of-thought, persona summaries, and preliminary fine-tuning. Full experimental details are provided in Technical Appendix \ref{sec:approach-details}. Overall, we find that the predictive accuracy of the answers simulated by the digital twins falls within a similar range across approaches (see Table \ref{tab:persona-accuracy-summary}). We hope this collection of baseline results will serve as a useful benchmark for future researchers exploring more advanced methods for persona training, such as reinforcement learning with human feedback (RLHF). For the initial analyses reported here, we focus on the text format using GPT4.1-mini.
\begin{table}[htbp]
    \centering
    \caption{Various persona simulation approaches and evaluation results}
    \begin{tabular}{c|c|c}
     Approach    & LLM &  Accuracy\\
     \hline
     Text Persona  & GPT4.1-mini & 71.72\% \\
  Text Persona & Gemini-flash2.5 & 69.40\% \\
     JSON Persona & GPT4.1-mini & 70.48\% \\
     JSON Persona & GPT4.1 & 71.05\% \\
     Persona Summary & GPT4.1-mini & 68.02\% \\
     Persona Summary + JSON Persona & GPT4.1-mini & 67.88\% \\
     Text Persona (Reasoning) & GPT4.1-mini & 70.39\%\\
     Text Persona (Repeating Questions) & GPT4.1-mini & 70.45\% \\
     Text Persona (Default Temperature) & GPT4.1-mini & 71.24\% \\
     JSON Persona (Predicted Output) & GPT4.1-mini & 69.00\% \\
     JSON Persona (Predicted Output) & GPT4.1 & 71.92\%\\
     LLM Finetuning (500 training samples) & GPT4.1-mini & 69.61\% \\
     Random Guessing & -- & 59.17\% \\
     Human Test/Retest & -- & 81.72\%
\end{tabular}
\label{tab:persona-accuracy-summary}
\end{table}


Figure \ref{fig:accuracy} reports, for each task, the predictive accuracy of the answers simulated by the digital twins, as well as the accuracy of a random benchmark which chooses each answer from a random uniform distribution. On average, across the 17 tasks the accuracy of the digital twin predictions is 71.72\%, and the ratio of the digital twin accuracy to the test-retest accuracy is 87.67\%. The improvement over the baseline is consistent across all question types, highlighting the value of personalization and LLM-based simulation.




\begin{figure}[htbp]
  \centering
  \includegraphics[width=\textwidth]{figs/accuracy_dist.png}
  \caption{Predictive accuracy}
  \label{fig:accuracy}
\end{figure}

\subsection{Aggregate level}
We test whether the data simulated from the digital twins replicates the average treatment effects from the 11 classic between-subject studies and the 5 classic within-subject studies included in our experiment. Table \ref{tab:replication} shows that for 6 of the 10 results replicated by waves 1-3 and wave 4, the results from the digital twins also replicate the results. For anchoring and adjustment, the digital twins replicate the effect when asking participants to estimate the height of the highest redwood tree. But when asking participants to estimate the number of African countries in the UN, 98.8\% of the twins gave the correct answer (54) and no anchoring effect was found. Three other between-subject effects were not replicated. In the outcome bias experiment, participants evaluate a physician's decision to operate on a patient. Humans evaluate the decision more favorably when the operation succeeded than when it failed, despite the risk being greater in the first condition. In contrast, digital twins all gave a favorable rating (``correct'' to ``clearly correct''), with no significant different across conditions. In the sunk cost fallacy experiment, the effect was actually reversed with the digital twins vs. their human counterparts, which we hope future research can explore. In the Allais problem experiment, which tests for violation of the independence axiom of utility theory, all digital twins chose the lower risk - lower reward option over the higher risk - higher reward one. Humans, on the other hand, were much more split in their decisions, and showed systematic differences across conditions (which violates the independence axiom of utility theory).  Finally, the base rate fallacy, which was replicated neither in wave 2 nor in wave 4, was not replicated by the digital twins either.

Moving to within-subject experiments, we find that the digital twin results match the human results in two of the five within-subject experiments (see Table \ref{tab:replication}). In the nonseparability of risk and benefit judgments study, the digital twins judgments display negative correlation as predicted, but the correlation is significant for only one of the items. For probability matching vs. maximizing, the digital twins always selected the normative option while their human counterparts chose the normative option about 30\% of the time only. For omission bias, participants were asked whether they would accept a vaccine that prevents catching a flu that has a 10\% chance of killing affected patients when the vaccine itself carries a 5\% chance of death. While approximately 45\%  of the human participants (45.10\% in wave 2, 44.80\% in wave 4) refused the vaccine, only 4.0\% of the twins refused the vaccine. This finding echoes our finding related to outcome bias where digital twins were much more favorable to medical professionals compared to their human counterparts.

Finally, we construct average demand curves from the pricing study. Figure \ref{fig:demand} shows the average demand curves from the responses from waves 3 vs. wave 4 vs. digital twins. We see that the average demand curves from wave 3 vs. 4 are practically indistinguishable. We find that the average demand curve obtained from the twin is not fully downward sloping, due to the twins' responses to free products. This echoes \cite{gui2023challenge}, although our digital twins produce demand curves that are downward sloping for positive prices and that are generally closer to the ground truth compared to the demand curves obtained by \cite{gui2023challenge} without such input data.

In sum, while digital twins replicate most between-subject and within-subject effects, there are notable exceptions. Some occur when digital twins fail to mimic the suboptimal or non-normative behaviors of humans, or cannot ``unlearn'' certain facts (e.g., the number of African countries in the UN). In some areas, twins do match suboptimal human responses (e.g., absolute vs.\ relative savings, less is more, dominator neglect), raising the broader question of whether digital twins should be seen as ``improved'' humans or as models that also replicate human deviations from normative behavior and knowledge gaps.

Other deviations appear in the medical domain (e.g., outcome bias, omission bias). Future research should examine whether digital twins are systematically more trusting of the medical profession than humans. Another factor may be that certain topics, such as vaccination, have become highly polarized, making it difficult for GPT models to reflect conservative viewpoints \citep{motoki2024more}. For instance, in our false consensus task, while about 45\% of humans somewhat or strongly supported increased deportations, 74.1\% of digital twins strongly or somewhat \emph{opposed} the measure. More research is needed to systematically study where digital twins diverge from humans, especially in medical and political domains, and to identify other domains where such differences may arise.


\tiny
    \begin{longtable}{p{2.8cm}p{2.8cm}p{4.8cm}
        >{\centering\arraybackslash}p{1.2cm}>{\centering\arraybackslash}p{1.2cm}>{\centering\arraybackslash}p{1.2cm}}
                    \caption{Replications of heuristics and biases}
        \label{tab:replication}\\
        \toprule
        \textbf{Task} & \textbf{Source} & \textbf{Prediction} & \multicolumn{3}{c}{\textbf{Replicated}} \\
        \multicolumn{3}{l}{}&Waves 1-3&Wave 4&Twins \\
                \midrule
       \multicolumn{6}{l}{\emph{Between-subject experiments}} \\
                       \midrule
   Base rate problem&\cite{kahneman1973psychology}&	no difference in prob. assessment when base rate=30 vs. 70	& \ding{55}&\ding{55}&\ding{55}\\
         \midrule
         Outcome bias&\cite{baron1988outcome}&average correctness assessment higher in success vs. failure condition&	\checkmark &\checkmark&\ding{55}\\
         \midrule
         Sunk cost fallacy&\cite{stanovich2008relative}&average number of purchases higher in sunk cost vs. no sunk cost condition &\checkmark&\checkmark&\ding{55}\\
        \midrule
        Allais problem&\cite{stanovich2008relative}&violation of independence axiom of utility theory (different choices in Form 1 vs. 2)	&\checkmark&\checkmark&\ding{55}\\
        \midrule
        Framing problem	&\cite{tversky1981framing}&	stronger preference for risky option under loss frame vs. gain frame	&\checkmark&\checkmark&\checkmark\\
        \midrule
        Conjunction problem (Linda) & \cite{tversky1983extensional}&	probability assessments higher for feminist bank teller vs. bank teller&\checkmark&\checkmark&\checkmark\\
        \midrule
        Anchoring and adjustment &\cite{tversky1974judgment,epley2004perspective}&average prediction higher with large vs. small anchor	&\checkmark \checkmark&\checkmark \checkmark&\checkmark \ding{55}\\
        \midrule
        Absolute vs. relative savings	&\cite{stanovich2008relative}&	probability of driving to store higher when discount is larger vs. smaller \% of price	&\checkmark&\checkmark&\checkmark\\
        \midrule
        Myside bias	&\cite{stanovich2008relative}&	average agreement higher for ban of German car in US vs. American car in Germany	&\checkmark&\checkmark&\checkmark\\
        \midrule
        Less is More&\cite{stanovich2008relative}	&average attractiveness higher when possibility of loss vs. no possibility of loss	&\checkmark&\checkmark&\checkmark\\
        \midrule
        WTA/WTP – Thaler problem&\cite{stanovich2008relative}&WTA-certainty$>$WTP-certainty$>$WTP-noncertainty&	\checkmark & \checkmark&\checkmark \\
        \bottomrule

       \multicolumn{6}{l}{\emph{Within-subject experiments}} \\
    \midrule

        False consensus	&\cite{furnas2024people}&	overpredict (underpredict) public support if own support (oppose)	&\checkmark&\checkmark& \checkmark\\
        \midrule
        Nonseparability of risk and benefits judgments&\cite{stanovich2008relative}&	negative correlation between benefits and risks for each item	&\checkmark \checkmark \checkmark \ding{55}& \checkmark \checkmark \checkmark \ding{55}&\checkmark \ding{55} \ding{55} \ding{55}\\
        \midrule
        Omission bias&\cite{stanovich2008relative}&	significant proportion avoid treatment&	\checkmark&	\checkmark&\ding{55}\\
        \midrule
        Probability matching vs. maximizing	&\cite{stanovich2008relative}&	significant proportion choose suboptimal strategy	&\checkmark&\checkmark&\ding{55}\\
        \midrule
        Dominator neglect &\cite{stanovich2008relative}&	significant proportion choose non-normative option&	\checkmark&	\checkmark&\checkmark\\
  \bottomrule
\end{longtable}
\normalsize


\section{Conclusion}
We present a unique dataset spanning over 500 questions and 2,000 respondents, with high data quality evidenced by strong correlations, good test-retest accuracy, and replication of known effects. While this resource can benefit social scientists broadly, our primary focus is on using it to build digital twins. These twins predict human behavior with out-of-sample accuracy reaching 87\% of the test-retest benchmark. Replication of average treatment effects is generally good, though further research is needed to determine if digital twins can capture non-normative behaviors and reflect the full diversity of political and domain-specific views. The dataset’s focus on the US and social science topics is a potential limitation. Overall, we hope this resource accelerates LLM research and social science applications while being mindful of societal risks such as dehumanization of research and excessive reliance on AI in decision-making.