EconBase
← All papers

Econometrics with Pre-Trained Embeddings for Unstructured Data

Yuya Shimizu

arXiv 19 Jul 2026 · Econometrics

arXiv:2607.17378 · PDF · Extracted main text

Abstract

Unstructured data, such as images and text, are increasingly used in empirical economics. Since training machine-learning models on unstructured data is costly, economists often use off-the-shelf pre-trained deep learning models developed by computer scientists to extract embeddings, which are then used as covariates in target economic analyses. Despite the popularity of this practice, its theoretical foundations remain limited. There are two main difficulties. First, the pre-trained model is usually trained on a different dataset and for a different task. Consequently, it is unclear when such a model can be used reliably for the target task. Second, the embedding function is subject to an identification problem, which makes it difficult to analyze the estimation error of the embedding function and its effect on the target task. In this paper, we provide sufficient conditions to overcome these difficulties and derive the convergence rate of machine learning models with pre-trained embeddings. We illustrate the theory through double machine learning applications for estimating parameters of interest, such as partially linear regression with unstructured controls, price elasticity in demand estimation considering the product quality measured by images and text, missing data imputation with unstructured data, and the average treatment effect with unstructured confounders.

Citation extraction

65
references
118
in-text mentions
65
distinct cited
0
self-citations
18,391
main-text words

appendix boundary found by appendix_command · 39% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Chernozhukov, Victor and Chetverikov, Denis and Demirer, Mert and Du… (2018) Double/debiased machine learning for treatment and structural parameters0.87452100%
2Watkins, Austin and Ullah, Enayat and Nguyen-Tang, Thanh and Arora,… (2023) Optimistic rates for multi-task representation learning0.8435360%
3P. Bajari and Z. Cen and V. Chernozhukov and M. Manukonda and S. Vij… (2025) Hedonic prices and quality adjusted price indices powered by AI0.84333100%
4Tripuraneni, Nilesh and Jordan, Michael and Jin, Chi (2020) On the theory of transfer learning: The importance of task diversity0.8307457%
5Farrell, Max H and Liang, Tengyuan and Misra, Sanjog (2021) Deep neural networks for estimation and inference0.7946450%
6Bach, Philipp and Chernozhukov, Victor and Klaassen, Sven and Spindl… (2024) Adventures in demand analysis using AI0.73732100%
7Schulte, Rickmer and Rügamer, David and Nagler, Thomas (2025) Adjustment for confounding using pre-trained representations0.73732100%
8Devlin, Jacob and Chang, Ming-Wei and Lee, Kenton and Toutanova, Kri… (2019) Bert: Pre-training of deep bidirectional transformers for language understanding0.6443267%
9Compiani, Giovanni and Morozov, Ilya and Seiler, Stephan (2025) Demand estimation with text and image data0.64422100%
10van der Vaart, AW and Wellner, Jon A (2023) Weak Convergence and Empirical Processes: With Applications to Statistics0.6069422%

Showing the top 10 of 65 scored citations.