EconBase
← All papers

Decision Theory for the Archetype Discovery Problem

José Luis Montiel Olea, Amilcar Velez, Zhuoheng Xu, Haomin Yu, Shunqi Zhang

arXiv 12 Jun 2026 · Econometrics

arXiv:2606.15002 · PDF · DOI · OpenAlex · Extracted main text

Abstract

In the archetype discovery problem a researcher wants to summarize N heterogeneous policy effects of interest that vary over a discrete set of covariates. The goal is to partition the set of covariates into K<N groups -- the archetype sets -- and to provide a summary of the policy effects for each group. We use decision theory to show that, under a weighted mean-squared-error criterion, a procedure analogous to the Sorted Group Average Treatment Effects (GATES) solves the archetype discovery problem. The key difference is that, in the optimal procedure, archetype sets are obtained by weighted K-means clustering of the N heterogeneous policy effects, instead of relying on K equally-spaced quantiles. We show that the procedure that minimizes average risk for a given prior can be obtained by clustering the different values of the posterior mean estimate of the policy effects of interest. Similarly, an approximately minimax procedure in large samples can be obtained by clustering a consistent estimator of the policy effects. In both of these cases, an exact solution to the weighted K-means clustering problem can be found using a simple and well-known dynamic programming algorithm.

Citation extraction

46
references
108
in-text mentions
46
distinct cited
0
self-citations
13,865
main-text words

appendix boundary found by appendix_command · 73% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Chernozhukov, V., M. Demirer, E. Duflo, and I. Fernández-Val (2025) b): Fisher–Schultz Lecture: Generic Machine Learning Inference on Heterogeneous Treatment Effects in Randomized Experiments, Wit…1.000185100%
2Breza, E., A. G. Chandrasekhar, and D. Viviano (2025) Generalizability with ignorance in mind: learning what we do (not) know for archetypes discovery1.000104100%
3Bruce, J. D (1965) Optimum quantization1.00064100%
4Lloyd, S (1982) Least squares quantization in PCM1.00053100%
5Wang, H. and M. Song (2011) Ckmeans0.92843100%
6Wu, X. and J. Rokne (1989) An O (KN lg N) algorithm for optimum K-level quantization on histograms of N points, in0.87462100%
7Grnlund, A., K. G. Larsen, A. Mathiasen, J. S. Nielsen, S. Schneider… (2017) Fast exact k-means, k-medians and Bregman divergence clustering in 1D0.81142100%
8Kim, K., J. Kim, and E. H. Kennedy (2026) Causal k-means clustering0.81142100%
9Wu, X (1991) Optimal quantization by matrix searching0.73732100%
10MacQueen, J. B (1967) Some methods of classification and analysis of multivariate observations, in0.64422100%

Showing the top 10 of 46 scored citations.