EconBase
← All papers

Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach

Jacob Carlson

arXiv 3 Nov 2025 · Econometrics

arXiv:2511.01680 · PDF · Extracted main text

Abstract

Social scientists are increasingly turning to unstructured datasets to unlock new empirical insights, e.g., estimating descriptive statistics of or causal effects on quantitative measures derived from text, audio, or video data. In many such settings, unsupervised analysis is of primary interest, in that the researcher does not want to (or cannot) pre-specify all important aspects of the unstructured data to measure; they are interested in "discovery." This paper proposes a general and flexible framework for pursuing discovery from unstructured data in a statistically principled way. The framework leverages recent methods from the literature on machine learning interpretability to map unstructured data points to high-dimensional, sparse, and interpretable "dictionaries" of concepts; computes statistics of dictionary entries for testing relevant concept-level hypotheses; performs selective inference on these hypotheses using algorithms validated by new results in high-dimensional central limit theory, producing a selected set ("discoveries"); and both generates and evaluates human-interpretable natural language descriptions of these discoveries. The proposed framework has few researcher degrees of freedom, is fully replicable, and is cheap to implement -- both in terms of financial cost and researcher time. Applications to recent descriptive and causal analyses of unstructured data in empirical economics are explored. An open source Jupyter notebook is provided for researchers to implement the framework in their own projects.

Citation extraction

68
references
238
in-text mentions
68
distinct cited
1
self-citations
17,894
main-text words

appendix boundary found by appendix_titled_section at “Appendix” · 52% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Bricken, Trenton and Templeton, Adly and Batson, Joshua and Chen, Br… (2023) Towards Monosemanticity: Decomposing Language Models With Dictionary Learning1.00083100%
2Movva, Rajiv and Peng, Kenny and Garg, Nikhil and Kleinberg, Jon and… (2025) Sparse Autoencoders for Hypothesis Generation1.00063100%
3Ludwig, Jens and Mullainathan, Sendhil (2024) Machine Learning as a Tool for Hypothesis Generation1.00053100%
4Bursztyn, Leonardo and Egorov, Georgy and Haaland, Ingar and Rao, Aa… (2023) Justifying Dissent0.97614393%
5Stantcheva, Stefanie (2024) Why Do We Dislike Inflation?0.96510390%
6Modarressi, Iman and Spiess, Jann and Venugopal, Amar (2025) Causal Inference on Outcomes Learned from Text0.9507386%
7Chernozhuokov, Victor and Chetverikov, Denis and Kato, Kengo and Koi… (2022) Improved central limit theorem and bootstrap approximations in high dimensions0.9416383%
8Belloni, Alexandre and Chernozhukov, Victor and Chetverikov, Denis a… (2018) High-Dimensional Econometrics and Regularized GMM0.92815580%
9Jiang, Nick and Sun, Xiaoqing and Dunlap, Lisa and Smith, Lewis and… (2025) Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit0.92843100%
10Paulo, Gonçalo and Mallen, Alex and Juang, Caden and Belrose, Nora (2024) Automatically Interpreting Millions of Features in Large Language Models0.9098475%

Showing the top 10 of 68 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Econometrics with Pre-Trained Embeddings for Unstructured Data0.40511