Gordon Burtch, Edward McFowland III, Mochen Yang, Gediminas Adomavicius
arXiv 6 Mar 2023 · Econometrics
arXiv:2303.02820 · PDF · Extracted main text
Despite increasing popularity in empirical studies, the integration of machine learning generated variables into regression models for statistical inference suffers from the measurement error problem, which can bias estimation and threaten the validity of inferences. In this paper, we develop a novel approach to alleviate associated estimation biases. Our proposed approach, EnsembleIV, creates valid and strong instrumental variables from weak learners in an ensemble model, and uses them to obtain consistent estimates that are robust against the measurement error problem. Our empirical evaluations, using both synthetic and real-world datasets, show that EnsembleIV can effectively reduce estimation biases across several common regression specifications, and can be combined with modern deep learning techniques when dealing with unstructured data.
appendix boundary found by appendix_command · 68% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Yang, M., McFowland III, E., Burtch, G., and Adomavicius, G (2022) Achieving reliable causal inference with data-mined variables: A random forest approach to the measurement error problem self | 1.000 | 10 | 4 | 100% |
| 2 | Fong, C. and Tyler, M (2021) Machine learning predictions as regression covariates | 1.000 | 5 | 4 | 100% |
| 3 | Allon, G., Chen, D., Jiang, Z., and Zhang, D (2023) Machine learning and prediction errors in causal inference | 0.928 | 4 | 3 | 100% |
| 4 | Nevo, A. and Rosen, A. M (2012) Identification with imperfect instruments | 0.928 | 4 | 3 | 100% |
| 5 | Yang, M., Adomavicius, G., Burtch, G., and Ren, Y (2018) Mind the gap: Accounting for measurement error and misclassification in variables generated via data mining self | 0.811 | 4 | 2 | 100% |
| 6 | Küchenhoff, H., Mwalili, S. M., and Lesaffre, E (2006) A general method for dealing with misclassification in regression: The misclassification SIMEX | 0.737 | 3 | 2 | 100% |
| 7 | Stefanski, A. L. A. and Cook, J. R (1995) Simulation-Extrapolation : The Measurement Error Jackknife | 0.737 | 3 | 2 | 100% |
| 8 | Qiao, M. and Huang, K.-W (2021) Correcting misclassification bias in regression models with variables generated via data mining | 0.737 | 3 | 2 | 100% |
| 9 | Wei, Y. and Malik, N (2022) Unstructured data, econometric models, and estimation bias | 0.737 | 3 | 2 | 100% |
| 10 | Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C… (2018) Double/debiased machine learning for treatment and structural parameters: Double/debiased machine learning | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 60 scored citations.