arXiv 3 Nov 2025 · Econometrics
arXiv:2511.01680 · PDF · Extracted main text
Social scientists are increasingly turning to unstructured datasets to unlock new empirical insights, e.g., estimating descriptive statistics of or causal effects on quantitative measures derived from text, audio, or video data. In many such settings, unsupervised analysis is of primary interest, in that the researcher does not want to (or cannot) pre-specify all important aspects of the unstructured data to measure; they are interested in "discovery." This paper proposes a general and flexible framework for pursuing discovery from unstructured data in a statistically principled way. The framework leverages recent methods from the literature on machine learning interpretability to map unstructured data points to high-dimensional, sparse, and interpretable "dictionaries" of concepts; computes statistics of dictionary entries for testing relevant concept-level hypotheses; performs selective inference on these hypotheses using algorithms validated by new results in high-dimensional central limit theory, producing a selected set ("discoveries"); and both generates and evaluates human-interpretable natural language descriptions of these discoveries. The proposed framework has few researcher degrees of freedom, is fully replicable, and is cheap to implement -- both in terms of financial cost and researcher time. Applications to recent descriptive and causal analyses of unstructured data in empirical economics are explored. An open source Jupyter notebook is provided for researchers to implement the framework in their own projects.
appendix boundary found by appendix_titled_section at “Appendix” · 52% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Bricken, Trenton and Templeton, Adly and Batson, Joshua and Chen, Br… (2023) Towards Monosemanticity: Decomposing Language Models With Dictionary Learning | 1.000 | 8 | 3 | 100% |
| 2 | Movva, Rajiv and Peng, Kenny and Garg, Nikhil and Kleinberg, Jon and… (2025) Sparse Autoencoders for Hypothesis Generation | 1.000 | 6 | 3 | 100% |
| 3 | Ludwig, Jens and Mullainathan, Sendhil (2024) Machine Learning as a Tool for Hypothesis Generation | 1.000 | 5 | 3 | 100% |
| 4 | Bursztyn, Leonardo and Egorov, Georgy and Haaland, Ingar and Rao, Aa… (2023) Justifying Dissent | 0.976 | 14 | 3 | 93% |
| 5 | Stantcheva, Stefanie (2024) Why Do We Dislike Inflation? | 0.965 | 10 | 3 | 90% |
| 6 | Modarressi, Iman and Spiess, Jann and Venugopal, Amar (2025) Causal Inference on Outcomes Learned from Text | 0.950 | 7 | 3 | 86% |
| 7 | Chernozhuokov, Victor and Chetverikov, Denis and Kato, Kengo and Koi… (2022) Improved central limit theorem and bootstrap approximations in high dimensions | 0.941 | 6 | 3 | 83% |
| 8 | Belloni, Alexandre and Chernozhukov, Victor and Chetverikov, Denis a… (2018) High-Dimensional Econometrics and Regularized GMM | 0.928 | 15 | 5 | 80% |
| 9 | Jiang, Nick and Sun, Xiaoqing and Dunlap, Lisa and Smith, Lewis and… (2025) Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit | 0.928 | 4 | 3 | 100% |
| 10 | Paulo, Gonçalo and Mallen, Alex and Juang, Caden and Belrose, Nora (2024) Automatically Interpreting Millions of Features in Large Language Models | 0.909 | 8 | 4 | 75% |
Showing the top 10 of 68 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Econometrics with Pre-Trained Embeddings for Unstructured Data | 0.405 | 1 | 1 |