EconBase
← All papers

Pragmatic DML with AI-Learned Representations

Andres Aradillas Fernandez, Victor Chernozhukov, Carlos Cinelli, Sven Klaassen, Whitney Newey, Martin Spindler, Jan Teichert-Kluge, Suhas Vijaykumar

arXiv 1 Oct 2026 · Econometrics

arXiv:2610.01935 · PDF · Extracted main text

Abstract

Text, images, and other rich covariates are increasingly compressed into AI-learned representations and then used as controls in causal analysis. We study when this approach is valid and develop a practical framework for causal inference with learned representations. For a broad class of estimands, an imperfect representation distorts the target causal parameter by the product of two representation errors: one in the outcome regression and one in the balancing weight (or Riesz representer). This yields three constructive results. First, cross-fitted double machine learning (DML) provides valid Wald inference for the representation-dependent target. When representation errors are small, the same interval covers the causal parameter, and it can even attain the semiparametric efficiency bound. Second, fold-wise representation learning (or fine-tuning) is compatible with DML inference for the causal parameter. To this end, we develop convex- and star-aggregation pipelines for learning and combining representations. Third, when representation errors are substantial, we can provide interpretable sensitivity regions and root-$n$ inference for their endpoints. In a multi-modal demand application, seven representation-specific estimates and their star aggregate all imply a negative near-unit elasticity for rank-based price response, and the result remains robust over the reported sensitivity grid.

Citation extraction

25
references
55
in-text mentions
25
distinct cited
7
self-citations
19,336
main-text words

appendix boundary found by appendix_command · 53% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Chernozhukov, Victor and Newey, Whitney K. and Singh, Rahul (2022) Automatic Debiased Machine Learning of Causal and Structural Effects self1.00083100%
2Chernozhukov, Victor and Cinelli, Carlos and Newey, Whitney K. and S… (2026) Long Story Short: Omitted Variable Bias in Causal Machine Learning self1.00073100%
3Chernozhukov, Victor and Chetverikov, Denis and Demirer, Mert and Du… (2018) Double/Debiased Machine Learning for Treatment and Structural Parameters self0.81142100%
4van der Vaart, Aad W (1998) Asymptotic Statistics0.7374350%
5Bickel, Peter J. and Klaassen, Chris A. J. and Ritov, Ya'acov and We… (1993) Efficient and Adaptive Estimation for Semiparametric Models0.5112250%
6Chernozhukov, Victor and Newey, Whitney K. and Quintas-Martinez, Vic… (2024) Automatic Debiased Machine Learning via Riesz Regression self0.51121100%
7van der Laan, Mark J. and Polley, Eric C. and Hubbard, Alan E (2007) Super Learner0.51121100%
8Athey, Susan and Imbens, Guido W (2019) Machine Learning Methods That Economists Should Know About0.40511100%
9Audibert, Jean-Yves (2007) Progressive mixture rules are deviation suboptimal0.40511100%
10Belloni, Alexandre and Chernozhukov, Victor and Hansen, Christian (2014) High-Dimensional Methods and Inference on Structural and Treatment Effects self0.40511100%

Showing the top 10 of 25 scored citations.