George Gui, Seungwoo Kim
arXiv 30 Sep 2025 · Econometrics
arXiv:2509.25709 · PDF · Extracted main text
Pre-experiment stratification, or blocking, is a well-established technique for designing more efficient experiments and increasing the precision of the experimental estimates. However, when researchers have access to many covariates at the experiment design stage, they often face challenges in effectively selecting or weighting covariates when creating their strata. This paper proposes a Generative Stratification procedure that leverages Large Language Models (LLMs) to synthesize high-dimensional covariate data to improve experimental design. We demonstrate the value of this approach by applying it to a set of experiments and find that our method would have reduced the variance of the treatment effect estimate by 10%-50% compared to simple randomization in our empirical applications. When combined with other standard stratification methods, it can be used to further improve the efficiency. Our results demonstrate that LLM-based simulation is a practical and easy-to-implement way to improve experimental design in covariate-rich settings.
appendix boundary found by appendix_command · 96% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Bai, Y (2022) Optimality of matched-pair designs in randomized controlled trials | 1.000 | 10 | 4 | 100% |
| 2 | Greevy, R., Lu, B., Silber, J. H., and Rosenbaum, P (2004) Optimal multivariate matching before randomization | 0.928 | 4 | 4 | 100% |
| 3 | Athey, S. and Imbens, G. W (2017) The econometrics of randomized experiments | 0.928 | 4 | 3 | 100% |
| 4 | Bruhn, M. and McKenzie, D (2009) In pursuit of balance: Randomization in practice in development field experiments | 0.644 | 2 | 2 | 100% |
| 5 | Morgan, K. L. and Rubin, D. B (2012) Rerandomization to improve covariate balance in experiments | 0.644 | 2 | 2 | 100% |
| 6 | Barrera-Osorio, F., Linden, L. L., and Saavedra, J. E (2019) Medium- and long-term educational consequences of alternative conditional cash transfer designs: Experimental evidence from colo… | 0.585 | 3 | 1 | 100% |
| 7 | de Mel, S., McKenzie, D., and Woodruff, C (2019) Labor drops: Experimental evidence on the return to additional labor in microenterprises | 0.585 | 3 | 1 | 100% |
| 8 | Abel, M., Burger, R., and Piraino, P (2020) The value of reference letters: Experimental evidence from south africa | 0.511 | 2 | 1 | 100% |
| 9 | Gerber, A., Hoffman, M., Morgan, J., and Raymond, C (2020) One in a million: Field experiments on perceived closeness of the election and voter turnout | 0.511 | 2 | 1 | 100% |
| 10 | Aher, G. V., Arriaga, R. I., and Kalai, A. T (2023) Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 33 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | LLM Personas as a Substitute for Field Experiments in Method Benchmarking | 0.405 | 1 | 1 |
| 2 | AI-Assisted Variance Reduction in Randomized Experiments | 0.405 | 1 | 1 |