EconBase
← All papers

LLM Personas as a Substitute for Field Experiments in Method Benchmarking

Enoch Hyunwook Kang

arXiv 24 Dec 2025 · Artificial Intelligence

arXiv:2512.21080 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Field experiments (A/B tests) are often the most credible benchmark for methods (algorithms) in societal systems, but their cost and latency bottleneck rapid methodological progress. LLM-based persona simulation offers a cheap synthetic alternative, yet it is unclear whether replacing humans with personas preserves the benchmark interface that adaptive methods optimize against. We prove an if-and-only-if characterization: when (i) methods observe only the aggregate outcome (aggregate-only observation) and (ii) evaluation depends only on the submitted artifact and not on the method's identity or provenance (method-blind evaluation), swapping humans for personas is just panel change from the method's point of view, indistinguishable from changing the evaluation population (e.g., New York to Jakarta). Furthermore, we move from validity to usefulness: we define an information-theoretic discriminability of the induced aggregate channel and show that making persona benchmarking as decision-relevant as a field experiment is fundamentally a sample-size question, yielding explicit bounds on the number of independent persona evaluations required to reliably distinguish meaningfully different methods at a chosen resolution.

Citation extraction

42
references
59
in-text mentions
42
distinct cited
1
self-citations
6,840
main-text words

appendix boundary found by appendix_command · 55% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Kang, Enoch Hyunwook and Yoganarasimhan, Hema (2025) TextBO: Bayesian Optimization in Language Space for Eval-Efficient Self-Improving AI self0.7218338%
2Blum, Avrim and Hardt, Moritz (2015) The ladder: A reliable leaderboard for machine learning competitions0.64422100%
3Bogert, Eric and Lauharatanahirun, Nina and Schecter, Aaron (2022) Human preferences toward algorithmic advice in a word association task0.5112250%
4Chiang, Wei-Lin and Zheng, Lianmin and Sheng, Ying and Angelopoulos,… (2024) Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference0.5112250%
5de Quidt, Jonathan and Haushofer, Johannes and Roth, Christopher (2018) Measuring and Bounding Experimenter Demand0.5112250%
6Dominguez-Olmedo, Ricardo and Hardt, Moritz and Mendler-Dünner, Cele… (2024) Questioning the Survey Responses of Large Language Models0.5112250%
7Kohavi, Ron and Tang, Diane and Xu, Ya (2020) Trustworthy online controlled experiments: A practical guide to a/b testing0.5112250%
8Levitt, Steven D. and List, John A (2011) Was There Really a Hawthorne Effect at the Hawthorne Plant? An Analysis of the Original Illumination Experiments0.5112250%
9Osborne, Merrick R. and Bailey, Erica R (2025) Me vs. the machine? Subjective evaluations of human- and AI-generated advice0.5112250%
10Toubia, Olivier and Gui, George Z and Peng, Tianyi and Merlau, Danie… (2025) Database report: Twin-2k-500: A data set for building digital twins of over 2,000 people based on their answers to over 500 ques…0.5112250%

Showing the top 10 of 42 scored citations.