EconBase
← All papers

Debiased Machine Learning for Unobserved Heterogeneity: High-Dimensional Panels and Measurement Error Models

Facundo Argañaraz, Juan Carlos Escanciano

arXiv 18 Jul 2025 · Econometrics

arXiv:2507.13788 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Developing robust inference for models with nonparametric Unobserved Heterogeneity (UH) is both important and challenging. We propose novel Debiased Machine Learning (DML) procedures for valid inference on functionals of UH, allowing for partial identification of multivariate target and high-dimensional nuisance parameters. Our main contribution is a full characterization of all relevant Neyman-orthogonal moments in models with nonparametric UH, where relevance means informativeness about the parameter of interest. Under additional support conditions, orthogonal moments are globally robust to the distribution of the UH. They may still involve other high-dimensional nuisance parameters, but their local robustness reduces regularization bias and enables valid DML inference. We apply these results to: (i) common parameters, average marginal effects, and variances of UH in panel data models with high-dimensional controls; (ii) moments of the common factor in the Kotlarski model with a factor loading; and (iii) smooth functionals of teacher value-added. Monte Carlo simulations show substantial efficiency gains from using efficient orthogonal moments relative to ad-hoc choices. We illustrate the practical value of our approach by showing that existing estimates of the average and variance effects of maternal smoking on child birth weight are robust.

Citation extraction

134
references
364
in-text mentions
149
distinct cited
5
self-citations
52,937
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Bonhomme, Stéphane (2012) Functional Differencing1.000259100%
2Chamberlain, Gary (1992) Efficiency bounds for semiparametric regression1.000146100%
3Belloni, Alexandre, Victor Chernozhukov, Christian Hansen, and Damia… (2016) Inference in high-dimensional panel models with an application to gun control1.000134100%
4Chernozhukov, Victor, Juan Carlos Escanciano, Hidehiko Ichimura, Whi… (2022) Locally robust semiparametric estimation self1.000133100%
5Arellano, Manuel and Stéphane Bonhomme (2012) Identifying distributional characteristics in random coefficients panel data models1.000126100%
6Honoré, Bo E and Martin Weidner (2024) Moment conditions for dynamic panel logit models with fixed effects1.00097100%
7Chen, Xiaohong and Andres Santos (2018) Overidentification in regular models1.00093100%
8Carrasco, Marine, Jean-Pierre Florens, and Eric Renault (2007) Chapter 77 Linear Inverse Problems in Structural Econometrics Estimation Based on Spectral Decomposition and Regularization1.00073100%
9Luenberger, David G (1997) Optimization by vector space methods1.00063100%
10Argañaraz, Facundo and Juan Carlos Escanciano (2023) On the Existence and Information of Orthogonal Moments For Inference self1.00054100%

Showing the top 10 of 149 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1xtdml: Double Machine Learning Estimation to Static Panel Data Models with Fixed Effects in R0.40511
2Double Machine Learning for Static Panel Data with Instrumental Variables: New Method and Applications0.40511